← get-fableCONTENT HISTORY

Update to get-fable

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.5.1

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Drive testable behavior changes and bug fixes through disciplined red-green-refactor cycles with observable regression tests. Use when implementing new features with unit/integration tests, fixing reproducible bugs, modifying business logic, or writing test-first behavior contracts — even if the user does not explicitly say \"fable-tdd\" (e.g. \"fix this bug test-first\", \"write a test and make it pass\", \"add this feature with tests\", \"TDD this logic\"). Do NOT use for broad exploratory prototyping without clear assertions (use fable-discover or fable-plan) or for post-implementation reviews (use fable-review).",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 425
    },
    {
      "relative_path": "evals/scenarios.json",
      "size_in_bytes": 4651
    },
    {
      "relative_path": "examples/failing-test-first.md",
      "size_in_bytes": 387
    },
    {
      "relative_path": "references/red-green-refactor.md",
      "size_in_bytes": 1084
    },
    {
      "relative_path": "references/test-strategy-and-hard-cases.md",
      "size_in_bytes": 4417
    },
    {
      "relative_path": "skill.package.json",
      "size_in_bytes": 436
    },
    {
      "relative_path": "templates/tdd-cycle.template.md",
      "size_in_bytes": 707
    }
  ],
  "name": "fable-tdd",
  "skill_md_contents": "---\nname: fable-tdd\ndescription: \"Drive testable behavior changes and bug fixes through disciplined red-green-refactor cycles with observable regression tests. Use when implementing new features with unit/integration tests, fixing reproducible bugs, modifying business logic, or writing test-first behavior contracts — even if the user does not explicitly say \\\"fable-tdd\\\" (e.g. \\\"fix this bug test-first\\\", \\\"write a test and make it pass\\\", \\\"add this feature with tests\\\", \\\"TDD this logic\\\"). Do NOT use for broad exploratory prototyping without clear assertions (use fable-discover or fable-plan) or for post-implementation reviews (use fable-review).\"\nversion: 1.3.0\npack: build\ninputs:\n  - behavior_contract\nrequires:\n  - test_harness\nproduces:\n  - regression_test\n  - behavior_change\ngates:\n  - red_observed\n  - green_observed\nfallback: fable-recover\nmutatesWorkspace: true\nparallelSafe: false\nneural_links:\n  precursors:\n    - fable-plan\n  continuations:\n    - fable-execute\n    - fable-verify\n  lateral_peers:\n    - fable-execute\n  recovery: fable-recover\n---\n\n# Fable TDD\n\nProve the behavior is missing or broken before changing production code, then make the smallest change that satisfies the right test at the right level.\n\n## Mission\nTDD is not \"write any failing test first.\" The red state must demonstrate the intended behavior gap through a trustworthy harness. A syntax error, broken fixture, stale build, or mock-only expectation does not earn the right to change production code.\n\nThe Skill optimizes for three things:\n- **causal confidence**: the test fails because the behavior is wrong;\n- **minimal intervention**: implementation changes only what the behavior requires;\n- **durable regression proof**: the test would catch the bug if it returned.\n\n## Activate When\n- fixing a reproducible bug or regression;\n- adding behavior with a stable enough contract to assert;\n- changing validation, state transitions, calculations, API behavior, persistence, or integration semantics;\n- refactoring where an invariant needs executable characterization first.\n\n## Do Not Activate When\n- the behavior cannot yet be located or reproduced (`fable-discover`);\n- the main uncertainty is external API semantics (`fable-research`);\n- the change is non-executable docs/metadata with no meaningful behavior test (`fable-execute`);\n- the required test harness is itself broken or executing stale artifacts (`fable-recover`).\n\n## Change Classification\nClassify the behavior before choosing a test.\n\n| Class | Preferred proof |\n| --- | --- |\n| Pure/domain logic | focused unit/property test |\n| Boundary validation/error mapping | unit or contract test at boundary |\n| Cross-module interaction | integration test through the changed contract |\n| Database/queue/cache behavior | integration test with realistic boundary where feasible |\n| HTTP/CLI/public API | contract/integration test at public entry point |\n| UI user flow | component/integration first; E2E for high-value cross-boundary behavior |\n| Concurrency/timing | deterministic coordination test, not sleep-and-hope |\n| Legacy behavior with poor seams | characterization test at nearest stable boundary |\n| Refactor/no intended behavior change | characterization/invariant tests before movement |\n\nUse the lowest test level that proves the real behavior **without mocking away the thing under test**.\n\n## Protocol\n\n### Stage 1 — Write the behavior contract\nState:\n- initial state/input;\n- trigger/action;\n- expected observable result;\n- relevant side effects;\n- error/boundary behavior;\n- invariant that must remain true.\n\nFor a bug, capture the concrete reproduction separately from the proposed implementation.\n\n### Stage 2 — Validate the harness\nBefore RED, confirm the chosen test actually executes the relevant path.\n\nCheck when applicable:\n- source vs built artifact;\n- test discovery/config;\n- fixture realism;\n- feature flags/env;\n- mock boundaries;\n- asynchronous completion;\n- cleanup/isolation;\n- whether the assertion observes public behavior rather than an internal call count.\n\n### Stage 3 — Choose the test level\nAsk:\n1. What production change would make this test fail again?\n2. Does the test cross the contract where the bug actually lives?\n3. Have I mocked the suspected failure away?\n4. Can a narrower test prove the same user-visible behavior more deterministically?\n\nIf you cannot answer #1, the test is probably weak.\n\n### Stage 4 — RED\nWrite the smallest test that captures the behavior contract and run it.\n\nRED is valid only if:\n- test executes;\n- failure is deterministic enough to reason about;\n- failure message/state matches the expected missing/broken behavior;\n- it is not failing because of setup, syntax, import, environment, stale artifact, or unrelated test pollution.\n\nIf the failure reason is wrong, repair the test/harness and repeat RED. **Do not touch production code yet.**\n\n### Stage 5 — Minimal GREEN\nChange the smallest production surface that can satisfy the valid RED.\n\nDo not:\n- redesign adjacent APIs;\n- add speculative flexibility;\n- refactor unrelated code;\n- weaken the assertion;\n- replace real behavior with a mock to reach green.\n\nRun the focused test until GREEN.\n\n### Stage 6 — Adjacent falsification\nBefore refactor, probe the closest failure surfaces appropriate to the change:\n- boundaries/empty/invalid input;\n- error path;\n- state transition ordering;\n- idempotency/retry;\n- concurrency/race;\n- compatibility with prior format/API;\n- persistence/transaction behavior.\n\nAdd tests only when they protect a meaningful contract; do not inflate count mechanically.\n\n### Stage 7 — Refactor under green\nNow improve structure if needed. Refactor in small steps and rerun the affected tests after each meaningful mutation.\n\n### Stage 8 — Fresh handoff\nRecord final mutation and hand off to `fable-verify` with:\n- behavior contract;\n- exact RED evidence/reason;\n- GREEN command/result;\n- tests added/changed;\n- production surfaces changed;\n- adjacent cases probed;\n- residual risks not covered by the focused test.\n\n## Decision Rules\n- Test passes before production change → it does not prove the requested gap; redesign the test or confirm behavior already exists.\n- Test errors before assertion → fix harness/test, not production.\n- Bug is nondeterministic → first control or instrument the nondeterminism; repeated random reruns are not a reliable RED.\n- Concurrency bug → prefer barriers/latches/fake clocks/deterministic scheduling over arbitrary sleeps.\n- Legacy code has no unit seam → test the nearest stable public boundary before introducing a seam; do not perform a broad refactor just to make unit testing aesthetically pure.\n- External dependency cannot be exercised locally → use a contract/fake only after establishing what behavior the fake must preserve from primary evidence.\n- User asks to \"just patch it\" → if executable behavior is changing, preserve RED discipline unless the absence of a viable harness is explicit and accepted.\n- Test expectation conflicts with current agreed product contract → do not force implementation to an obsolete test; resolve the contract first.\n- A test only asserts that a mock was called → add observable outcome/state evidence unless the call itself is the public contract.\n\n## Invariants\n- No production behavior mutation before a valid RED for testable changes.\n- RED and GREEN refer to the same behavior contract.\n- Tests are not weakened to make implementation pass.\n- The suspected failure mechanism is not mocked away.\n- Final GREEN is fresh after the last relevant mutation.\n- Refactoring does not add behavior outside the accepted card.\n\n## Failure Taxonomy\n### Wrong RED\nFailure comes from syntax/import/setup/fixture/environment rather than target behavior. Repair harness first.\n\n### False GREEN\nTest passes but does not cross the real failure boundary, often because a mock or stale artifact bypasses it. Strengthen/reposition the test.\n\n### Flaky RED/GREEN\nOutcome changes without relevant code mutation. Identify nondeterminism before treating either state as evidence.\n\n### Untestable legacy surface\nNo narrow seam exists. Characterize at a stable boundary, then introduce the smallest seam supported by the test.\n\n### Contract ambiguity\nExpected behavior itself is disputed/unclear. Return to planning/research rather than encoding a guess as a test.\n\n### Implementation loop\nValid RED exists but two materially similar fixes fail. Stop editing and route to `fable-recover` with the RED, attempts, and observed failure differences.\n\n## Anti-Patterns\n- writing implementation, then backfilling a test that immediately passes;\n- accepting any failure as RED;\n- changing expected output to match current implementation;\n- asserting only mock interactions when user-visible state can be asserted;\n- using sleeps to \"test\" a race;\n- building an elaborate test abstraction before proving one behavior;\n- broad legacy refactor before characterization;\n- forcing unit tests where only an integration boundary can prove the behavior;\n- skipping fresh GREEN after refactor;\n- equating coverage percentage with regression proof.\n\n## Evidence Packet\n\n```text\nBehavior contract:\nTest level + why:\nRED command/result + expected failure reason:\nProduction mutation:\nGREEN command/result:\nAdjacent cases probed:\nRefactor mutations:\nFinal fresh GREEN:\nResidual risk / verify next:\n```\n\n## Completion Criteria\nTDD completes when:\n- the behavior gap was demonstrated by a valid RED;\n- production code changed only after that RED;\n- focused GREEN is observed for the same contract;\n- meaningful adjacent failure surfaces were considered;\n- final tests are fresh after refactor/last mutation;\n- `fable-verify` receives enough evidence to independently falsify the result.\n\n## Progressive Resources\n- Deep strategy: `references/test-strategy-and-hard-cases.md`\n- Existing cycle reference: `references/red-green-refactor.md`\n- Example: `examples/failing-test-first.md`\n"
}

SHA-256 of public snapshot: ad4aa1b00884ac7bf10c0b3fcb30362967b3f9badb180de9bf0bc1d53db47e00