get-fable
Mamdouh Abo Ammar v1.5.1
Publisher description
From the marketplace listing
Routes software work through 25 focused lifecycle skills across 8 packs, keeps durable mutation-aware state in .fable/state.json, and requires fresh evidence before completion.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Skill instructions
fable-artifact8.72 KB
---
name: fable-artifact
description: "Design and author structured technical proposals, responsive artifacts, architecture diagrams, Mermaid charts, and interactive components. Use when creating standalone markdown reports, architectural specifications, Mermaid diagrams, or interactive artifact widgets — even if the user does not explicitly say \"fable-artifact\" (e.g. \"create an artifact\", \"draw an architecture diagram\", \"write a technical proposal\", \"generate a mermaid flowchart\"). Do NOT use for quick one-line conversational answers."
version: 1.3.0
pack: system
inputs:
- artifact_spec
requires:
- design_requirements
produces:
- artifact_document
- interactive_ui
gates:
- hierarchy_clear
- theme_adaptive
fallback: fable-plan
mutatesWorkspace: true
parallelSafe: true
neural_links:
precursors:
- fable-dataviz
- fable-plan
continuations:
- fable-verify
- fable-review
lateral_peers:
- fable-dataviz
recovery: fable-recover
---
# Fable Artifact
Turn requirements and evidence into a standalone artifact whose content, structure, diagrams, links, and rendered output can all be checked independently.
## Mission
An artifact is not successful because the Markdown parses or the diagram looks polished. It must communicate the right claims to the right audience, preserve source truth, make assumptions visible, and render in the medium the user will actually consume.
## Activate When
- producing RFCs, architecture proposals, technical reports, runbooks, decision records, diagrams, or rich standalone documentation;
- packaging analysis/results into a durable document;
- creating a diagram or interactive explanatory artifact from established requirements/evidence.
## Do Not Activate When
- a one-paragraph answer is sufficient;
- application logic itself needs modification (`fable-execute`);
- the main work is numerical chart selection (`fable-dataviz`);
- facts required by the artifact have not been researched/discovered yet.
## Artifact Classification
| Artifact | Primary contract |
| --- | --- |
| RFC/proposal | decision, alternatives, trade-offs, acceptance/rollout |
| Architecture doc | boundaries, data/control flow, invariants, deployment |
| Runbook | trigger, prerequisites, executable steps, rollback/escalation |
| Incident/report | timeline/evidence/impact without invented causality |
| Decision record | context, decision, alternatives, consequences |
| Diagram | semantic topology/sequence/state, not decorative boxes |
| Interactive explainer | accessible interaction + deterministic content/source |
## Protocol
### Stage 1 — Define audience and job
State who will use the artifact and what decision/action it must support. One artifact can contain multiple sections, but it should have one primary communication job.
### Stage 2 — Build a claim/source map
Before writing polished prose, identify:
- measured/source-backed facts;
- design decisions;
- assumptions/inferences;
- unresolved items;
- data/visual sources;
- claims that need citations/links.
Do not invent examples, metrics, architecture components, test results, or citations to make the document feel complete.
### Stage 3 — Choose structure from artifact type
Examples:
- RFC: context → goals/non-goals → evidence → options → decision → design → risks → rollout/rollback → acceptance;
- runbook: symptom/trigger → safety prerequisites → diagnosis → action → verification → rollback/escalation;
- architecture: scope → context → components/boundaries → flows → data/state → failure/security → deployment/operations → decisions.
Avoid generic template sections that do not serve the document's job.
### Stage 4 — Design diagrams semantically
Each node/edge should represent a real component, state, dependency, event, or data/control flow. Label direction/meaning where ambiguity exists.
For Mermaid or other text diagrams:
- validate syntax;
- quote/escape labels safely;
- avoid giant unreadable graphs;
- split views by question (context, container, sequence, state) when one diagram becomes overloaded.
### Stage 5 — Write for scanning and decision quality
Use hierarchy, short sections, tables only when comparison benefits, callouts sparingly, and explicit `Decision`, `Risk`, `Unknown`, `Evidence` language where useful.
Do not turn every sentence into bullets or bury the main decision below background detail.
### Stage 6 — Validate internal consistency
Cross-check:
- terms/names/versions match throughout;
- diagram matches prose;
- examples match actual API/schema;
- links/anchors exist;
- recommendations follow cited evidence;
- no section contradicts the accepted design.
### Stage 7 — Render in the target medium
A source file is not the final artifact when rendering matters. Preview/render Markdown/Mermaid/HTML/PDF/other target as appropriate and inspect for clipping, broken diagrams, overflow, missing assets, unreadable typography, or inaccessible interactions.
### Stage 8 — Run truth and usability review
Ask:
- can the audience locate the primary decision/action quickly?
- which statements are fact vs proposal vs inference?
- are any claims unsupported?
- would a diagram imply a relationship that does not exist?
- can a future reader execute/maintain the artifact without chat context?
## Decision Rules
- Prefer explicit `unknown/not checked` to filling gaps with plausible content.
- A diagram should answer a question; split it when one view mixes topology, sequence, deployment, and ownership beyond readability.
- Use links/citations for external/current claims where traceability matters; do not fabricate sources.
- Generated charts/tables inherit their data provenance from DataViz/research packets.
- If artifact changes as the design changes, reconcile every diagram/example/decision reference before completion.
- Interactive artifacts need keyboard/accessibility/fallback behavior appropriate to their use; visual novelty is not a substitute for communication.
- Keep implementation details out of an executive/decision artifact unless they materially change the decision.
- Use the repository/designated artifact destination; do not litter roots/temp paths.
## Invariants
- Every factual claim is source-backed or clearly labeled as inference/proposal.
- No fake data, citations, test results, architecture components, or case studies.
- Diagram/prose/examples describe the same design.
- Target rendering is checked when rendering is part of delivery.
- Artifact stands alone without requiring hidden conversation context.
- Audience can identify the primary decision/action.
## Failure Taxonomy
### Content hallucination
Missing fact is filled with plausible detail. Remove/research/label unknown.
### Diagram semantic drift
Diagram no longer matches current design or implies false flow. Regenerate/reconcile.
### Template bloat
Generic sections obscure the document's job. Remove sections without decision value.
### Source/render mismatch
Markdown/HTML code looks valid but target renderer clips/breaks assets/diagram. Verify rendered output.
### Internal contradiction
Versions/names/decisions differ across sections. Establish one canonical claim map and reconcile.
### Audience mismatch
Document is technically complete but too low/high altitude for intended reader. Reorganize around their decision/action.
### Link/source rot
Critical evidence points to invalid/missing paths. Repair or preserve relevant content locally when appropriate.
## Anti-Patterns
- fake metrics to make proposal persuasive;
- Mermaid diagram full of decorative boxes with no clear edge semantics;
- generic "Overview / Architecture / Conclusion" template regardless of job;
- copying raw research notes without synthesis;
- using callout blocks everywhere;
- claiming rendered/accessible without previewing target medium;
- architecture prose updated while diagram remains stale;
- citations that do not support the sentence;
- giant wall of implementation detail before the main decision.
## Artifact Packet
```text
Audience / job:
Artifact type + destination:
Claim/source map:
Decisions vs assumptions vs unknowns:
Structure rationale:
Diagram(s) + question answered:
Data/source provenance:
Internal consistency checks:
Rendered validation:
Accessibility/usability check:
Known limitations:
```
## Completion Criteria
Artifact completes when:
- content is evidence-grounded and audience/job-aligned;
- structure supports the intended decision/action;
- diagrams/examples/data are semantically consistent;
- unknowns are explicit rather than invented;
- target rendering/links/assets are validated;
- a zero-context reader can understand and use the artifact.
## Progressive Resources
- Deep guide: `references/source-grounded-artifact-design.md`
- Existing guide: `references/artifact-composition-guide.md`
- Example: `examples/system-architecture-diagram.md`
Referenced files: 7
fable-config7.37 KB
---
name: fable-config
description: "Configure and audit AI agent harness settings, permissions allowlists, environment variables, editor keybindings, and lifecycle hook integrations. Use when modifying settings.json, adjusting tool permissions, setting up environment variables, or configuring agent lifecycle hooks — even if the user does not explicitly say \"fable-config\" (e.g. \"update settings\", \"allow command permissions\", \"configure agent hooks\", \"setup environment variables\"). Do NOT use for application-level business configuration."
version: 1.3.0
pack: system
inputs:
- config_change
requires:
- target_harness
produces:
- settings_diff
gates:
- valid_json
- safe_permissions
fallback: fable-plan
mutatesWorkspace: true
parallelSafe: false
neural_links:
precursors:
- fable-plan
continuations:
- fable-verify
- get-fable
lateral_peers:
- fable-plan
recovery: fable-recover
---
# Fable Config
Change harness/configuration behavior with explicit precedence, least privilege, secret-safe handling, and a verified rollback path.
## Mission
Configuration is executable behavior. A syntactically valid JSON/TOML/YAML file can still disable a guard, broaden permissions, write to the wrong scope, or be ignored because another source has higher precedence.
The Skill must prove both **configuration validity** and **effective behavior**.
## Activate When
- changing host/harness settings, permissions, hooks, keybindings, model/tool config, or environment references;
- adding/removing lifecycle enforcement;
- troubleshooting why a setting is not taking effect;
- configuring safe command/tool allowlists;
- migrating configuration formats/scopes.
## Do Not Activate When
- storing raw credentials/secrets;
- application business logic is the real change;
- a host capability is unknown and needs research/discovery first;
- the requested change is to weaken a safety boundary merely to bypass an error without understanding it.
## Configuration Classification
| Change | Main risk |
| --- | --- |
| Permission/allowlist | excessive privilege/wildcard scope |
| Hook registration | hook exists but is never invoked / blocks wrong phase |
| Environment config | precedence, secret leakage, type/format |
| Host settings | wrong user/project scope, unsupported keys |
| Model/tool config | changed behavior/cost/access unexpectedly |
| Keybindings/UI | collision/override |
| Generated config | manual edit overwritten by generator |
## Protocol
### Stage 1 — Locate source and effective scope
Identify:
- target host/version;
- project vs user/global scope;
- all configuration sources and precedence;
- generated vs hand-maintained files;
- existing user customizations to preserve.
Do not edit the first file with the right name until you know it is effective.
### Stage 2 — Define desired behavior and least privilege
State exactly what capability should become allowed/blocked/triggered and what must remain unchanged.
For permissions, prefer narrow command/tool/path patterns. Wildcards need explicit justification and threat consideration.
### Stage 3 — Protect secrets
Configuration may reference environment variable names or secure stores, but should not embed raw tokens/passwords/private keys unless the target's secure format explicitly requires protected encrypted storage.
If an existing secret is found in plaintext, do not echo it; route exposure handling appropriately.
### Stage 4 — Make a minimal merge
Preserve unknown/user-defined settings. Avoid replacing an entire config object/file when one scoped key can be merged.
For generated config, modify source/generator then regenerate.
### Stage 5 — Validate structure and semantics
Run available parser/schema/host diagnostics. Check:
- syntax;
- key types/enums;
- duplicate/conflicting entries;
- unsupported/deprecated keys;
- permission pattern scope;
- hook command/path existence.
### Stage 6 — Prove effective behavior
A valid file is not enough. Test the configured behavior safely:
- target command is allowed while broader command remains denied;
- hook fires at intended lifecycle point;
- project override wins as expected;
- setting is visible to the target host;
- rollback restores previous behavior.
### Stage 7 — Record rollback and handoff
Capture changed source, effective scope, validation, behavioral proof, and how to revert if the host fails after restart/update.
## Decision Rules
- Preserve existing user settings not owned by the task.
- Prefer exact allowlists over `*`, shell wildcards, root/global permissions.
- A config key accepted by parser but ignored by host is not a successful change; verify effect.
- Host integration levels differ; do not claim hooks/enforcement where the host only supports advisory rules.
- Environment variable name may be stored; secret value should remain in secure environment/credential store.
- If config precedence is uncertain, investigate before editing more files.
- Do not disable security/approval checks simply because they are blocking an unsafe action.
- If a malformed edit can lock out the host, keep a reversible backup/atomic write strategy.
## Invariants
- Least privilege is preserved or any broadening is explicit and justified.
- Raw secrets are not introduced into repository/config logs.
- Existing unrelated user configuration is preserved.
- Edited source is the effective source-of-truth.
- Syntax/schema and actual host behavior both verify.
- Rollback is possible for material changes.
## Failure Taxonomy
### Precedence mismatch
Edited file is shadowed by another scope/source. Trace effective configuration.
### Schema-valid but ignored
Host accepts file but does not support/use key. Verify host version/capability.
### Permission overreach
Pattern allows more than requested. Narrow and test negative case.
### Hook misregistration
Script exists but lifecycle never invokes it. Validate registration path/event and permissions.
### Secret leakage
Credential embedded/logged. Remove safely and rotate if exposure occurred.
### User-config clobber
Whole-file rewrite loses existing settings. Restore/merge surgically.
### Generated drift
Manual edit is overwritten. Change generator/source instead.
## Anti-Patterns
- `"allow": ["*"]` to stop approval prompts;
- editing global config when project scope suffices;
- copying API tokens into settings examples;
- validating JSON syntax and calling the hook configured;
- claiming full lifecycle enforcement on a host with advisory-only integration;
- resetting the whole config file for one key;
- disabling security controls to make a command pass;
- manually patching generated config.
## Configuration Packet
```text
Target host/version/scope:
Effective config sources + precedence:
Desired behavior:
Settings changed:
Permissions before/after:
Secret handling:
Syntax/schema validation:
Behavioral proof + negative case:
Preserved user settings:
Rollback:
```
## Completion Criteria
Configuration completes when:
- correct effective source/scope was changed minimally;
- configuration parses and satisfies schema/capability constraints;
- permissions remain least-privilege;
- secrets/unrelated settings remain safe;
- target behavior is empirically observed, including important negative case;
- rollback and host limitations are explicit.
## Progressive Resources
- Deep guide: `references/config-precedence-permissions-and-hooks.md`
- Existing rules: `references/harness-configuration-rules.md`
- Example: `examples/configure-allowlist.md`
Referenced files: 7
fable-cowork8.59 KB
---
name: fable-cowork
description: "Execute autonomous multi-step cowork sessions with silent tool chaining, outcome-first progress reporting, and strict safety boundary enforcement. Use when executing complex background tasks autonomously, running multi-tool refactoring workflows without conversational noise, or performing deep automated passes — even if the user does not explicitly say \"fable-cowork\" (e.g. \"run autonomously in background\", \"cowork mode\", \"execute quietly and report results\", \"silent refactor pass\"). Do NOT use when active user dialogue or step-by-step confirmation is required."
version: 1.3.0
pack: system
inputs:
- autonomous_task
requires:
- scoped_goal
produces:
- autonomous_deliverable
- outcome_summary
gates:
- no_mid_chain_noise
- outcome_first
fallback: fable-execute
mutatesWorkspace: true
parallelSafe: true
neural_links:
precursors:
- get-fable
continuations:
- fable-execute
- fable-verify
- fable-handoff
lateral_peers:
- fable-spark
recovery: fable-recover
---
# Fable Cowork
Carry a scoped engineering objective through multiple tool calls and lifecycle stages with minimal conversational interruption, while preserving the same safety, evidence, and scope rules as interactive work.
## Mission
Autonomy is not permission to improvise indefinitely. Cowork mode should reduce conversational overhead—not remove checkpoints, verification, stop conditions, or user intent.
The agent owns execution between meaningful boundaries. It must still stop when the task requires a new product decision, destructive authorization, unavailable external permission, or a contradiction that changes the agreed goal.
## Activate When
- the user explicitly delegates a multi-step task end-to-end;
- the objective and success criteria are clear enough to continue without routine questions;
- the work may span discovery, planning, implementation, verification, review, and handoff;
- repeated narration would add noise while tool evidence can drive progress.
## Do Not Activate When
- the job is a simple answer or one bounded edit;
- core requirements are genuinely ambiguous and different choices materially change the product;
- an external irreversible action lacks authorization;
- work must wait for a future event rather than execute now;
- autonomy would require inventing credentials, access, or user intent.
## Autonomy Classification
| Task state | Cowork posture |
| --- | --- |
| Clear bounded sequence | execute silently through evidence gates |
| Multiple independent cards | delegate with explicit ownership |
| New architecture decision emerges | stop/replan; do not choose silently if material |
| Repeated failure | enter recovery before more mutation |
| External auth/permission unavailable | stop at hard wall with exact handoff |
| Destructive/public action authorized | execute with release/security checks |
| Destructive/public action not authorized | prepare/verify only; do not perform it |
## Protocol
### Stage 1 — Lock the objective and boundaries
Before long execution, capture:
- requested outcome;
- non-goals/protected surfaces;
- authorization boundaries;
- repository/worktree state;
- acceptance evidence;
- hard stop conditions.
Do not convert a broad aspiration into unlimited scope.
### Stage 2 — Route internally by lifecycle
Use the same specialist sequence as interactive get-fable work. Cowork is an execution mode across Skills, not a replacement for discovery/TDD/recovery/review.
### Stage 3 — Work in bounded chunks
For each card/phase:
- perform the necessary tool calls;
- collect evidence immediately;
- update state/checkpoint;
- continue only if the next dependency is already authorized and determined.
Keep user-visible updates sparse but meaningful for long work: surfaced finding, changed blocker, completed milestone—not narration of every command.
### Stage 4 — Detect drift
Stop autonomous forward motion if:
- requested outcome changed;
- new architecture/product choice has material trade-offs not implied by existing intent;
- scope expands substantially;
- user-owned/unrelated work conflicts with required mutation;
- security/destructive boundary appears;
- external authentication/permission cannot be obtained safely.
### Stage 5 — Handle failures through recovery
Do not hide failure inside silent loops. Repeated similar failures route to `fable-recover`; record hypothesis changes and only resume mutation when new evidence justifies it.
### Stage 6 — Verify before declaring outcome
Run fresh, relevant verification after final mutations. For substantial work include review/security/release gates as appropriate. If a required environment cannot be exercised, report INCOMPLETE rather than converting autonomy into fabricated confidence.
### Stage 7 — Preserve resumability
For long work or any hard wall, create a handoff/checkpoint containing exact state, evidence, blockers, and next safe action.
### Stage 8 — Report outcome first
Final response should lead with what actually happened, then evidence, then remaining blockers. Do not replay the internal tool sequence.
## Decision Rules
- Autonomy reduces questions only when existing intent is sufficient; it does not authorize material product decisions that were never implied.
- Prefer progress over clarification for recoverable implementation details; prefer stopping over guessing for irreversible/security/public-contract decisions.
- Do not silently reset/overwrite unrelated user changes to make the workspace convenient.
- Long task does not justify scope expansion; create additional cards only when they are necessary to the requested outcome.
- Background-like work must still happen in the current execution context; do not promise asynchronous future completion unless a scheduling mechanism actually exists.
- If a public publish/tag/delete/migration action is outside current authorization, prepare and verify it but stop before the irreversible step.
- A final "done" requires fresh evidence; partial completion should be reported precisely when a hard wall remains.
## Invariants
- User objective and protected scope remain stable or changes are surfaced.
- All mutations still obey specialist gates.
- No blind retry loops are hidden by silent execution.
- External/destructive authorization boundaries are respected.
- Unrelated user changes and credentials remain protected.
- Long work remains resumable after interruption.
- Final claims are evidence-backed and outcome-first.
## Failure Taxonomy
### Scope drift
Agent discovers adjacent opportunities and starts implementing them. Return to required outcome/cards.
### Silent uncertainty
Agent makes a material product choice to avoid interrupting. Stop/replan when the choice is not implied by intent.
### Silent failure loop
Repeated tool failures are hidden behind autonomy. Route to recovery and surface the blocker if diagnosis cannot progress.
### Authorization wall
Required publish/delete/prod/credential action is not authorized or available. Stop at prepared state with exact next action.
### Workspace collision
Unrelated user work overlaps required files. Preserve it; coordinate/narrow rather than overwrite.
### Evidence gap
Implementation finished but required browser/provider/CI/external evidence cannot run. Report partial/incomplete, not success.
## Anti-Patterns
- "keep going no matter what" as autonomy policy;
- narrating every tool call despite cowork mode;
- asking permission for every reversible implementation detail;
- silently making a new product/architecture choice with material trade-offs;
- retrying until something turns green;
- hiding incomplete external gates behind a polished final summary;
- expanding the task because there is still token/time budget;
- promising background work that is not actually scheduled/executing.
## Cowork Checkpoint
```text
Goal / non-goals:
Authorization boundaries:
Current phase/card:
Completed outcomes:
Fresh evidence:
Protected unrelated state:
New decisions made + basis:
Failures/recovery state:
Hard blockers:
Primary next safe action:
Can continue autonomously? yes/no + why
```
## Completion Criteria
Cowork completes when:
- requested outcome is delivered or a genuine hard wall is reached;
- lifecycle gates remain intact despite reduced narration;
- no material scope/authorization decision was guessed;
- final evidence is fresh and limitations explicit;
- task can be resumed from a durable checkpoint if incomplete;
- report begins with actual outcome, not process narration.
## Progressive Resources
- Deep guide: `references/autonomy-boundaries-and-checkpoints.md`
- Existing discipline: `references/silent-execution-discipline.md`
- Example: `examples/silent-refactor-session.md`
Referenced files: 7
fable-dataviz9.19 KB
---
name: fable-dataviz
description: "Design and generate accessible, cohesive data visualizations, SVG charts, metric cards, and dashboard tiles with theme-adaptive styling and verified viewports. Use when creating SVG charts, rendering metrics plots, designing dashboard visuals, or visualizing performance trends — even if the user does not explicitly say \"fable-dataviz\" (e.g. \"make a chart of this data\", \"plot these benchmarks\", \"create an SVG graph\", \"visualize these metrics\"). Do NOT use for non-visual text-only data summaries or generic code edits."
version: 1.3.0
pack: system
inputs:
- data_source
requires:
- metric_specs
produces:
- visualization_artifact
- svg_chart
gates:
- theme_contrast_valid
- viewbox_defined
fallback: fable-execute
mutatesWorkspace: true
parallelSafe: true
neural_links:
precursors:
- fable-discover
continuations:
- fable-artifact
- fable-run
- fable-verify
lateral_peers:
- fable-artifact
recovery: fable-recover
---
# Fable DataViz
Turn data into a visual claim that is easy to read **without changing what the data actually says**.
## Mission
A chart is an argument about magnitude, trend, distribution, relationship, uncertainty, or composition. The first job is to choose a visual encoding that matches that question. The second is to preserve statistical meaning. Styling comes after both.
A valid SVG with attractive colors can still be a bad visualization if it truncates axes deceptively, aggregates incompatible groups, hides missing data, implies causality from correlation, or invents precision the source does not support.
## Activate When
- a metric/trend/distribution/comparison/relationship needs visual explanation;
- benchmark/eval results need charts or stat graphics;
- a report/artifact needs an evidence-backed visual;
- raw data must be transformed into SVG/chart code or a visual specification.
## Do Not Activate When
- there is no actual data and the request would require inventing values;
- a plain table is more accurate/readable for a small lookup task;
- the core task is document structure rather than visual encoding (`fable-artifact`);
- the user requests an illustrative image rather than a data visualization.
## Question Classification
| Question | Useful first encoding | Typical misuse |
| --- | --- | --- |
| Compare categories | bar/dot plot | pie with many similar slices |
| Trend over ordered time | line/area with careful baseline | unordered category line chart |
| Distribution | histogram/box/violin/dot | average-only bar |
| Relationship | scatter/bubble with scale caveats | dual-axis correlation theater |
| Part-to-whole | stacked bar/100% bar; limited pie | sum components that are not one whole |
| Ranking | sorted bar/dot | alphabetic order hiding rank |
| Single KPI + context | stat + baseline/change | giant number without denominator/timeframe |
| Uncertainty | interval/band/error bars | precise point with hidden variance |
## Protocol
### Stage 1 — Establish data provenance and semantic contract
Record:
- source/dataset/version/time window;
- unit and denominator;
- category/time definitions;
- missing/null semantics;
- whether values are counts, rates, percentages, currency, estimates, or modeled outputs;
- uncertainty/precision available.
If source or metric meaning is ambiguous, stop and resolve it before rendering.
### Stage 2 — State the visual question
Write one sentence: `This chart should help the reader see ___`.
If there are multiple unrelated questions, create separate views rather than forcing one overloaded chart.
### Stage 3 — Validate transformations
Before plotting, explicitly define:
- filters;
- grouping/aggregation;
- normalization/denominator;
- sorting;
- date bucketing/time zone;
- handling of missing/outliers;
- derived metrics/calculations.
Check totals/ranges before and after transformation. Never silently drop records that change the claim.
### Stage 4 — Choose encoding and scales
Use position/length for precise comparisons where possible. Choose linear/log/percentage scales based on metric semantics.
Baseline rules:
- bar length usually needs meaningful zero because length encodes magnitude;
- line/scatter axes may use non-zero domains if clearly labeled and not exaggerating the story;
- log scales require positive values and explicit labeling;
- dual axes are high-risk and need strong justification.
### Stage 5 — Encode uncertainty and data quality
If estimates have intervals/variance/sample sizes, show or state them when material. Mark missing periods/categories rather than connecting them as if observed.
Do not show more decimal places than source precision justifies.
### Stage 6 — Design for reading and accessibility
Prioritize:
- descriptive title stating metric/context;
- direct labels where they reduce legend decoding;
- readable typography/spacing;
- contrast and non-color cues;
- accessible title/description for SVG;
- responsive `viewBox`/appropriate container behavior;
- units and source note.
Do not rely on red/green or hue alone for meaning.
### Stage 7 — Validate the rendered artifact
Check:
- chart renders without clipping/overlap;
- data coordinates match source values;
- axes/ticks/labels/legend are correct;
- small/large screens when responsive;
- light/dark theme if required;
- accessibility metadata;
- no transformation/render code silently changes ordering or values.
### Stage 8 — Run a deception audit
Ask:
- would a reasonable reader infer a larger/smaller effect than raw data supports?
- is the denominator/time window obvious?
- are missing values hidden?
- does annotation imply causality not established?
- are categories incomparable due to different bases?
Fix the visual claim, not only the pixels.
## Decision Rules
- Never invent data, labels, sample sizes, sources, or benchmark results.
- A percentage without denominator/base often needs contextualization before visualization.
- Avoid pie/donut when readers need precise comparison or categories are numerous.
- Do not downsample by simply dropping points when extrema/events matter; use a documented aggregation/sampling strategy.
- Missing values are not zero unless the domain explicitly defines them that way.
- Sort categories to support the question unless natural/order semantics require otherwise.
- Use zero baseline for bars by default; exceptions require an encoding where truncation is not misleading and must remain visible.
- Correlation chart/temporal coincidence does not justify causal annotation.
- If a chart cannot remain legible at target size, simplify/segment rather than shrink labels into illegibility.
## Invariants
- Every plotted mark maps to source/transformation logic.
- Units, denominator, and timeframe remain truthful.
- Missing/uncertain data is not silently converted into certainty.
- Scale choices do not intentionally exaggerate magnitude.
- Accessibility does not depend on color alone.
- Rendered output can be traced back to source data and transformation steps.
## Failure Taxonomy
### Wrong chart question
Encoding answers composition while reader needs precise comparison. Re-select chart by analytical question.
### Aggregation distortion
Grouping/normalization changes denominator or hides important subgroup behavior. Recompute and document transformation.
### Scale deception
Axis/domain makes modest changes look extreme. Restore appropriate baseline/domain and labels.
### Missing-data fiction
Null periods are plotted as zero/interpolated without justification. Mark gaps or document imputation.
### Overplotting/crowding
Marks/labels overlap and conceal distribution. Aggregate, facet, sample responsibly, or change encoding.
### Accessibility failure
Contrast/color-only meaning/text size prevents interpretation. Add non-color cues/direct labels/accessible metadata.
### Source uncertainty
Metric meaning or provenance is unclear. Stop rendering and resolve source contract.
## Anti-Patterns
- starting with "make it a donut" before understanding the question;
- truncating bar axes to dramatize change;
- dual axes used to manufacture correlation;
- treating missing values as zero;
- inventing sample data to make the chart look complete;
- downsampling away spikes without disclosure;
- 3D/perspective effects that distort area/length;
- rainbow palettes with no semantic reason;
- title like "Revenue" with no unit/timeframe;
- declaring a chart verified because SVG syntax parses.
## Visualization Packet
```text
Question / intended takeaway:
Source + version/timeframe:
Metric definition / unit / denominator:
Transformations:
Missing/outlier/uncertainty handling:
Encoding + scale rationale:
Accessibility choices:
Rendered validation:
Deception audit:
Source note:
```
## Completion Criteria
Visualization completes when:
- visual question and metric semantics are explicit;
- transformations are reproducible and totals/ranges checked;
- encoding/scale match the analytical task without distortion;
- missing/uncertain data is honest;
- artifact renders accessibly at target size/theme;
- source/data-to-mark traceability exists;
- no claim exceeds what the data supports.
## Progressive Resources
- Deep guide: `references/truthful-chart-selection-and-validation.md`
- Existing design system: `references/dataviz-design-system.md`
- Example: `examples/render-bar-chart.md`
Referenced files: 7
fable-delegate7.61 KB
---
name: fable-delegate
description: "Delegate independent subtasks to parallel workers or subagents with strict disjoint ownership, bounded scope, and explicit acceptance contracts. Use when coordinating multiple independent tasks, parallelizing multi-file work without merge conflicts, or managing subagent task assignments — even if the user does not explicitly say \"fable-delegate\" (e.g. \"split this work among workers\", \"delegate these subtasks\", \"run parallel agents\", \"assign independent pieces\"). Do NOT use for tightly coupled sequential tasks or when shared mutable state causes edit collisions."
version: 1.3.0
pack: build
inputs:
- bounded_cards
requires:
- disjoint_ownership
produces:
- delegation_contracts
- worker_results
gates:
- ownership_explicit
- acceptance_explicit
fallback: fable-plan
mutatesWorkspace: true
parallelSafe: true
neural_links:
precursors:
- fable-plan
continuations:
- fable-execute
- fable-verify
lateral_peers:
- fable-tdd
recovery: fable-recover
---
# Fable Delegate
Use parallel workers when independence is real, not merely because there are multiple tasks.
## Mission
Delegation should reduce latency without multiplying ambiguity. The parent remains responsible for decomposition, contract stability, integration, and final proof.
A worker is not a place to dump context. It receives one bounded responsibility with enough evidence to act independently and a return contract that makes integration reviewable.
## Activate When
- two or more cards can proceed without waiting for each other's decisions;
- specialist perspectives can investigate independent hypotheses in parallel;
- separate components have stable interfaces and independent acceptance checks;
- independent review/research can be parallelized without shared mutation.
## Do Not Activate When
- workers would change the same unstable contract or invariant;
- one card's design result changes another card's requirements;
- a tiny task costs more to explain/integrate than to execute;
- the parent has not established enough context to write precise worker contracts;
- all workers depend on one unresolved architectural decision.
## Independence Classification
For each candidate pair, classify coupling:
| Coupling | Parallel? | Why |
| --- | --- | --- |
| Different files + different contracts | usually yes | low semantic overlap |
| Different files + shared unstable contract | no | semantic conflict despite path separation |
| Same integration file + separate components | maybe, but serialize integration | component work may parallelize; integration does not |
| Shared read-only evidence | yes | no ownership conflict |
| Same database schema/invariant | usually no | one worker can invalidate another's assumptions |
| Independent research hypotheses | yes | merge conclusions, not source edits |
| Parent decision required by both | no | resolve decision first |
## Delegation Protocol
### Stage 1 — Prove independence
Check:
- write ownership;
- shared contracts/invariants;
- dependency order;
- generated artifacts;
- migration/state coupling;
- integration hotspot;
- whether one worker failure changes another worker's assumptions.
If independence cannot be explained in one short paragraph, do not parallelize yet.
### Stage 2 — Write a worker contract
Each contract includes:
- objective;
- context/evidence already established;
- owned paths or semantic area;
- forbidden scope;
- dependencies assumed stable;
- acceptance evidence to produce;
- expected return packet;
- stop/escalation conditions.
Do not ask a worker to "handle X" with an entire repository and no boundary.
### Stage 3 — Decide mutation ownership
Prefer one writer per semantic area. Multiple read-only investigators may inspect overlapping files; multiple writers should not independently redefine the same contract.
For a shared integration file, assign integration to the parent or one named worker after component work converges.
### Stage 4 — Dispatch and supervise by exception
Do not micromanage every command. Intervene when:
- worker discovers a load-bearing unknown;
- owned scope must expand;
- acceptance cannot be satisfied;
- contract assumption is false;
- worker stalls/repeats failure.
### Stage 5 — Inspect returns, not summaries alone
Require concrete artifacts: diff/paths, commands/results, evidence, unresolved risk. A prose "done" is not integration evidence.
### Stage 6 — Integrate centrally
The parent checks:
- combined diff;
- contract compatibility;
- cross-worker behavior;
- generated artifacts;
- integration tests;
- whether worker-local evidence is still fresh after merge/integration mutations.
## Decision Rules
- Different files are neither necessary nor sufficient for independence.
- Stable contract + isolated implementation can parallelize; unstable contract + separate files should not.
- If workers produce competing approaches, keep them read-only until the parent chooses; do not merge both experiments blindly.
- If a worker needs to cross ownership boundaries, pause that worker and renegotiate the contract instead of silently expanding scope.
- A failed/hung worker does not automatically justify rerunning the same prompt; classify failure and route the card to `fable-recover` or execute locally.
- Parent integration verification is mandatory whenever delegated work affects one user-visible behavior.
## Invariants
- Every mutable surface has clear ownership.
- Workers receive only the context needed to make their bounded decisions.
- Worker completion never equals global completion.
- Shared contract changes have one integration owner.
- Combined workspace is verified after all integration mutations.
## Failure Taxonomy
### Ownership collision
Two workers need the same mutable contract/file. Stop parallel mutation and replan ownership.
### Semantic collision
Files differ but both change one invariant/schema. Treat as coupled work.
### Context starvation
Worker repeatedly asks questions the parent should have resolved. Improve contract or route back to discovery/plan.
### Scope escape
Worker finds necessary work outside ownership. Pause, evaluate, then expand/replan explicitly.
### Worker-local green / integration red
Local checks pass but combined behavior fails. Parent owns diagnosis; do not bounce workers blindly.
### Worker stall
No meaningful progress or repeated failure. Terminate/recover rather than consuming unbounded budget.
## Anti-Patterns
- "two tasks = two agents";
- path-disjointness as the only parallelism test;
- giving every worker the entire broad task;
- allowing each worker to update shared manifests/routes independently;
- accepting prose summaries without diffs/evidence;
- merging worker outputs before inspecting assumptions;
- assuming subagents transfer responsibility away from the parent;
- delegating a decision the parent has not made simply to avoid making it.
## Delegation Contract
```text
Worker objective:
Why independent:
Evidence/context:
Owns:
Must not change:
Stable dependencies assumed:
Acceptance evidence:
Stop/escalate if:
Return packet:
Integration owner:
```
## Completion Criteria
Delegation is complete when:
- every worker contract had explicit ownership and acceptance;
- returned artifacts/evidence were inspected;
- ownership/contract conflicts are resolved;
- combined diff passes integration verification on the current workspace;
- residual failures are routed explicitly rather than hidden in worker summaries.
## Progressive Resources
- Deep guide: `references/parallelism-and-integration.md`
- Existing contract guide: `references/subagent-contracts.md`
- Template: `templates/delegation-contract.template.md`
- Example: `examples/parallel-workers-walkthrough.md`
Referenced files: 7
fable-discover7.83 KB
---
name: fable-discover
description: "Gather the smallest set of repository, environment, documentation, and runtime evidence needed before planning or changing code. Use when load-bearing facts are unknown, tracing execution paths, inspecting unfamiliar packages, or resolving codebase contradictions — even if the user does not explicitly say \"fable-discover\" (e.g. \"explore the codebase\", \"how does this work\", \"where is this implemented\", \"find where this route is handled\"). Do NOT use when the bounded edit is already known (use fable-execute), for external API/version research (use fable-research), or for running existing test suites (use fable-verify)."
version: 1.3.0
pack: core
inputs:
- exploration_target
requires:
- codebase_access
produces:
- repository_evidence
- execution_path
gates:
- load_bearing_unknowns_resolved
fallback: fable-recover
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- get-fable
continuations:
- fable-research
- fable-plan
- fable-execute
lateral_peers:
- fable-memory
recovery: fable-recover
---
# Fable Discover
Build the smallest reliable mental model of an unfamiliar code path before anyone commits to a design or edit.
## Mission
Discovery is not "read a lot of files." It is uncertainty reduction.
The Skill should leave the next specialist with enough evidence to make a decision without rediscovering the repository, while stopping before exploration becomes archaeology.
## Activate When
- entry points, ownership, or execution flow are not known;
- a request spans an unfamiliar package/subsystem;
- generated code, plugins, runtime configuration, queues, jobs, or dependency injection may change the real execution path;
- the user describes behavior but not where it is implemented;
- a previous assumption about the repository has been contradicted.
## Do Not Activate When
- the exact bounded edit and target file are already known;
- the job is purely external API/version research (`fable-research`);
- the task is only to run existing checks (`fable-verify`).
## Situation Classification
Classify the unknown before searching.
| Unknown | First evidence to seek | Common trap |
| --- | --- | --- |
| Topology | manifests, workspace config, package boundaries | assuming root package owns runtime |
| Entry point | CLI/server/job/plugin bootstrap | starting from a similarly named helper |
| Execution path | calls, handlers, events, data transitions | following imports without proving runtime reachability |
| Configuration | env/schema/defaults/feature flags | reading defaults while production overrides them |
| Generated behavior | build scripts, codegen outputs, generated manifests | editing generated output instead of source |
| Plugin/extension path | registration, discovery, loader contracts | missing dynamic loading because grep finds no direct import |
| Persistence/data path | repositories, schemas, transactions, queues | stopping at service layer before side effects |
| Test architecture | harness, fixtures, test entry points | assuming tests exercise the same artifact users run |
## Discovery Protocol
### Stage 1 — Frame the unknowns
Write 2-7 questions whose answers would materially change the plan. Mark each as `load-bearing` or `nice-to-know`.
Examples:
- Which process receives this request?
- Is the package source executed directly or from `dist/`?
- Where is this plugin registered?
- Which storage boundary commits the state?
Do not start broad search until the questions are explicit.
### Stage 2 — Establish repository topology
Read project instructions and manifests first. Identify:
- workspace/package roots;
- build/test commands;
- generated directories;
- host/plugin manifests;
- source vs distribution entry points;
- configuration sources;
- relevant ownership boundaries.
### Stage 3 — Trace a real path
Start from an observable entry point and follow concrete symbols/events toward the target behavior.
For every hop record:
- file/symbol;
- why this hop is reachable;
- input/output contract;
- side effect or state transition;
- certainty: `[measured]`, `[inferred]`, or `[unresolved]`.
An import chain alone is not proof that code executes.
### Stage 4 — Probe runtime only when static evidence is insufficient
Use safe read-only runtime observation to answer a named question: logs, help output, route listing, test discovery, build metadata, or a narrow reproduction.
Do not mutate production state just to satisfy curiosity.
### Stage 5 — Resolve contradictions
When measured evidence contradicts the current model, update the model immediately. Do not preserve an old theory because it made the earlier search coherent.
### Stage 6 — Stop deliberately
Stop when every load-bearing question is either:
- answered with evidence; or
- explicitly unresolved with a named consequence and next Skill.
Nice-to-know questions do not block handoff.
## Decision Rules
- If the unknown is external and time-sensitive, hand that question to `fable-research` instead of inferring from model memory.
- If the exact edit becomes bounded and no design decision remains, hand off to `fable-execute`.
- If multiple components/risks must be coordinated, hand off to `fable-plan`.
- If search results suggest generated code, find the generator/source-of-truth before recommending edits.
- If direct imports disappear at a boundary, inspect registration tables, event buses, dependency injection, plugin discovery, reflection, codegen, and runtime configuration.
- If tests and runtime appear to disagree, record both paths; do not assume the test harness is authoritative.
## Invariants
- Discovery is read-only unless the user explicitly changes the task.
- Every load-bearing conclusion has concrete repository/runtime evidence.
- Inference is labeled as inference.
- Search breadth is justified by an unresolved question.
- Generated output is not treated as canonical source without proving it is hand-maintained.
## Failure Taxonomy
### Search miss
A symbol cannot be located. Check aliases, generated names, dynamic loading, registries, event dispatch, reflection, compiled output, and package boundaries.
### False path
Files look relevant but are not reachable from the real entry point. Return to a proven runtime/bootstrap boundary.
### Environment ambiguity
Behavior depends on env/flags/config. Identify precedence and which configuration is active; do not report a default as runtime fact.
### Source/artifact mismatch
Tests or commands execute built/stale artifacts rather than edited source. Record both paths and route to `fable-recover` if this caused repeated failure.
### Unknown remains load-bearing
Do not paper over it. Hand off to research or report the unresolved decision explicitly.
## Anti-Patterns
- reading every file in a package "for context";
- trusting filenames as architecture;
- treating grep frequency as importance;
- following imports without proving runtime reachability;
- ignoring build/codegen/plugin registration;
- reporting assumptions without `[inferred]` labels;
- continuing discovery after the planning decision is already safe.
## Evidence Packet / Handoff
Produce a compact packet:
```text
Target:
Load-bearing questions:
Measured facts: file:symbol → fact
Execution path: entry → ... → side effect
Configuration/runtime notes:
Generated/plugin boundaries:
Unresolved questions + consequence:
Recommended next Skill:
```
## Completion Criteria
Discovery is complete when the next specialist can explain:
- where execution starts;
- which path reaches the behavior;
- which contracts/state boundaries matter;
- which facts are measured vs inferred;
- what remains unknown;
- why those remaining unknowns do or do not block the next action.
## Progressive Resources
- Deep playbook: `references/repository-investigation-playbook.md`
- Existing evidence protocol: `references/evidence-gathering-protocol.md`
- Example: `examples/codebase-inspection.md`
Referenced files: 7
fable-eval8.56 KB
---
name: fable-eval
description: "Evaluate changes to agent prompts, skills, routing policies, and harnesses against reproducible baselines, held-out suites, and regression benchmarks. Use when optimizing agent system prompts, measuring skill triggering accuracy, evaluating routing changes, or running benchmark regressions — even if the user does not explicitly say \"fable-eval\" (e.g. \"benchmark this prompt\", \"evaluate agent behavior\", \"test skill performance\", \"run the eval suite\"). Do NOT use for routine application unit tests (use fable-verify)."
version: 1.3.0
pack: evolution
inputs:
- candidate_modification
requires:
- reproducible_baseline
produces:
- eval_verdict
- regression_evidence
gates:
- baseline_frozen
- holdout_tested
- rollback_defined
fallback: fable-plan
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- skill-creator
continuations:
- fable-plan
- fable-execute
lateral_peers:
- fable-verify
recovery: fable-recover
---
# Fable Eval
Measure whether an agent-control change improves the behavior it claims to improve without quietly overfitting the benchmark or breaking neighboring behavior.
## Mission
An eval is a decision instrument, not a scoreboard. It needs a frozen comparison point, representative semantic families, oracle isolation, explicit failure costs, and a rollback decision.
A candidate should not win because the prompts resemble its instructions, because holdouts leaked into authoring, or because one average score hides a severe regression.
## Activate When
- changing Skills, prompts, routers, hooks, agent profiles, policies, or model-control logic;
- comparing candidate prompt/agent configurations;
- measuring trigger precision/recall or action compliance;
- validating a new behavioral maturity claim;
- investigating whether an apparent improvement is robust or benchmark-specific.
## Do Not Activate When
- verifying ordinary application behavior (`fable-verify`);
- authoring a Skill before its intended behavior is clear (`skill-creator`);
- running a one-off subjective prompt demo with no acceptance decision.
## Evaluation Classification
| Change | Primary eval risk |
| --- | --- |
| Router/trigger | false positives, false negatives, precedence |
| Skill instruction | action correctness, forbidden shortcuts, boundary behavior |
| Spark/next-action | top-1 action, unsafe suggestion, silence precision |
| Hook/guard | enforcement, false blocking, bypasses |
| Prompt/persona | task quality + regressions + instruction conflicts |
| Tool policy | correct tool choice, unsafe/missing action |
| Model/config | quality/latency/cost variance across representative tasks |
## Protocol
### Stage 1 — Define the decision before running tests
State:
- candidate being evaluated;
- baseline/control;
- exact behavior expected to improve;
- metrics and thresholds;
- unacceptable regressions;
- rollback action.
Avoid inventing metrics after seeing results.
### Stage 2 — Build semantic scenario families
Cover distinct decisions, not wording variants. Include as applicable:
- straightforward positive case;
- non-trigger/boundary case;
- ambiguous competing action;
- adversarial shortcut pressure;
- partial/contradictory evidence;
- failure/recovery path;
- legacy/constrained environment;
- unseen holdout.
Record family coverage separately from raw prompt count.
### Stage 3 — Freeze baseline and corpus identity
Bind the run to:
- corpus hash/version;
- candidate/baseline identity;
- evaluator/scorer version;
- provider/model/config where external;
- timestamp/repository revision.
Changing the subject, oracle, corpus, or scoring logic invalidates direct comparability unless explicitly normalized.
### Stage 4 — Protect the oracle
Provider-facing requests must not reveal expected actions, forbidden actions, category labels, holdout identity, scoring implementation, or answer-bearing metadata.
Do not draft the candidate while repeatedly reading holdout failures. Promote discovered cases into a future checked corpus and preserve a new unseen holdout.
### Stage 5 — Execute baseline and candidate consistently
Use the same task inputs, tool availability, context budget, temperature/configuration, timeout policy, and scoring rules where comparison requires them.
Capture provider errors/timeouts as failures or explicit unavailable states; do not fill missing outputs from the oracle.
### Stage 6 — Score by slices, not average alone
Inspect:
- overall metric;
- each semantic family;
- negative/adversarial forbidden violations;
- high-cost regressions;
- variance/repeated-run stability where stochasticity is material;
- routing confusion pairs where applicable.
A 1% average gain is not acceptable if it introduces a severe release/security/recovery regression.
### Stage 7 — Investigate suspicious gains
Check for:
- prompt leakage;
- duplicated/near-duplicate scenarios;
- benchmark-specific keyword matching;
- changed tool/context budget;
- scorer drift;
- cherry-picked seeds/runs;
- examples copied into the candidate.
### Stage 8 — Decide and preserve rollback
Verdict:
- **ACCEPT**: thresholds met, no prohibited regression, evidence representative/fresh;
- **REJECT**: candidate regresses required behavior or fails threshold;
- **INCONCLUSIVE**: evidence lacks breadth/stability/holdout integrity.
Record baseline artifact so rollback remains possible.
## Decision Rules
- Semantic family breadth matters more than raw scenario count.
- Surface rewrites of one case do not create independent coverage.
- Holdouts stop being holdouts once used repeatedly to tune the candidate.
- Compare slices before averages; safety-critical forbidden violations can veto a higher average score.
- If stochastic variance could change the decision, repeat enough runs to estimate stability rather than cherry-picking one seed.
- If provider/runtime errors differ between candidate and baseline, separate infrastructure failure from behavior score.
- Never preserve an old maturity result after the evaluated Skill/corpus/oracle changes unless freshness validation proves identity.
- Do not lower thresholds after a candidate fails simply to ship it.
## Invariants
- Baseline is reproducible/frozen before candidate judgment.
- Oracle/holdout data remains hidden from the evaluated agent.
- Candidate and baseline are compared under equivalent conditions where claimed.
- Missing/failed provider outputs are never replaced by expected answers.
- Every acceptance decision has a rollback path.
- High-cost regressions remain visible even when aggregate score improves.
## Failure Taxonomy
### Benchmark overfit
Candidate improves checked cases but fails unseen family/holdout. Increase semantic breadth and reject/generalize candidate.
### Oracle leakage
Expected/forbidden/category data reaches provider/candidate authoring loop. Discard contaminated evidence and create fresh blind cases.
### Metric blindness
Average improves while important slice worsens. Use per-family and veto metrics.
### Non-comparable runs
Different model/tool/context/scorer settings produce apparent gain. Re-run under controlled conditions.
### High variance
Repeated runs change verdict. Increase samples/control nondeterminism or mark inconclusive.
### Corpus drift
Skill/scenario/oracle changed after evidence capture. Mark evidence stale and rerun.
## Anti-Patterns
- five paraphrases counted as five independent tests;
- reading holdouts while tuning every candidate;
- accepting on average score alone;
- changing thresholds after seeing failure;
- treating provider timeout as skipped rather than failed/incomplete;
- evaluating only positive examples;
- copying expected action vocabulary in a way that reveals the answer per case;
- claiming M4 because an evidence JSON file exists even though corpus hash changed.
## Eval Report
```text
Candidate / baseline:
Corpus + hashes:
Provider/config:
Semantic families:
Metrics + thresholds:
Per-family results:
Forbidden/high-cost regressions:
Variance/repeats:
Leakage/comparability checks:
Verdict: ACCEPT | REJECT | INCONCLUSIVE
Rollback:
Evidence freshness:
```
## Completion Criteria
Evaluation completes when:
- baseline and candidate identities are explicit;
- semantic coverage and blind holdout integrity are credible;
- metrics are inspected by meaningful slices;
- regressions/forbidden actions are not hidden by averages;
- verdict and rollback are evidence-backed;
- evidence freshness is tied to the exact evaluated corpus/control.
## Progressive Resources
- Deep guide: `references/benchmark-design-and-overfit-control.md`
- Existing protocol: `references/eval-harness-protocol.md`
- Example: `examples/eval-regression-run.md`
Referenced files: 7
fable-execute8.21 KB
---
name: fable-execute
description: "Implement one accepted, bounded work card with immediate local verification, invariant preservation, and zero scope drift. Use when executing a planned work card, applying a well-defined code change, implementing an isolated function, or performing targeted single-scope edits — even if the user does not explicitly say \"fable-execute\" (e.g. \"implement this card\", \"write the code for this step\", \"apply the agreed changes\", \"build this component\"). Do NOT use when new architectural decisions are required (use fable-plan) or when tests are repeatedly failing (use fable-recover)."
version: 1.3.0
pack: core
inputs:
- accepted_card
requires:
- bounded_scope
produces:
- implementation_diff
- card_completion
gates:
- invariants_preserved
- acceptance_checked
fallback: fable-plan
mutatesWorkspace: true
parallelSafe: false
neural_links:
precursors:
- fable-plan
- fable-tdd
- fable-delegate
continuations:
- fable-verify
lateral_peers:
- fable-simplify
recovery: fable-recover
---
# Fable Execute
Implement one accepted change without turning a bounded card into an unreviewable wandering session.
## Mission
Execution owns implementation, not architecture discovery. The agent should make the smallest coherent change that satisfies the accepted card, keep invariants visible, validate immediately, and stop when new facts invalidate the plan.
"Minimal" means minimal behaviorally complete diff, not necessarily the fewest changed lines.
## Activate When
- scope and acceptance are explicit;
- load-bearing architecture/external facts are settled;
- the card can be implemented without inventing a new public contract;
- the task is a bounded static/config/docs change where TDD adds no meaningful proof;
- a TDD Skill has already established RED and implementation is ready.
## Do Not Activate When
- a behavior change still needs a valid RED (`fable-tdd`);
- implementation reveals a new architecture decision (`fable-plan`);
- a repeated failure needs diagnosis (`fable-recover`);
- the task is to independently verify or review an existing diff (`fable-verify`/`fable-review`).
## Change Classification
Before editing, classify the card:
| Change | Execution concern |
| --- | --- |
| Local behavior | preserve surrounding contract |
| Config/manifest | precedence, generated source, packaging |
| Public API/CLI | compatibility and caller impact |
| Data/persistence | transaction/migration invariant |
| Generated output | edit source/generator, regenerate once |
| Dependency wiring | lifecycle/cleanup/error propagation |
| Repair after TDD | satisfy RED without widening behavior |
| Repair after review | fix only grounded finding + affected proof |
## Protocol
### Stage 1 — Re-read the card against current workspace
Confirm:
- objective;
- owned scope;
- prohibited scope;
- dependencies/assumptions;
- acceptance evidence;
- current git/workspace state.
If the workspace changed since planning in a way that invalidates assumptions, stop and replan rather than force the old card through.
### Stage 2 — Inspect before mutation
Read the exact target code and its immediate contracts/callers/tests. Avoid broad rediscovery, but do not edit from stale memory.
Identify invariants that must stay true during the change.
### Stage 3 — Choose the smallest coherent mutation
Prefer:
- existing repository patterns;
- source-of-truth over generated output;
- local compatibility over speculative abstractions;
- reversible changes when uncertainty remains;
- one behavior change at a time.
### Stage 4 — Mutate and record impact
Track:
- files/surfaces changed;
- new/changed contract;
- generated artifacts affected;
- acceptance command required after this mutation.
If an edit crosses card boundaries, pause and explain why before continuing.
### Stage 5 — Immediate acceptance check
Run the narrowest meaningful acceptance evidence as soon as the coherent mutation exists.
A command exit code alone is insufficient if it does not exercise the changed behavior.
### Stage 6 — Classify failure before another edit
On failure ask:
- expected RED from ongoing TDD?
- typo/compile issue introduced by current edit?
- incorrect implementation hypothesis?
- stale/build/config/harness problem?
- hidden dependency/architecture change?
Only make another mutation if the failure provides new information that justifies it.
### Stage 7 — Clean the diff
Before handoff:
- remove debug output/dead experiments;
- inspect accidental formatting or generated noise;
- verify no unrelated file changed;
- regenerate deterministic outputs from source when required;
- preserve user-owned concurrent changes.
### Stage 8 — Handoff for independent verification
Send `fable-verify`:
- card objective;
- diff/touched surfaces;
- acceptance command/result;
- invariants considered;
- generated/config impacts;
- residual risk and unverified surfaces.
## Decision Rules
- New load-bearing unknown → stop and route to discover/research/plan.
- New architecture/public contract decision → `fable-plan`; do not hide it as implementation detail.
- Behavior change with no valid regression test where one is feasible → `fable-tdd`.
- Generated file requested → locate generator/source and mutate source; regenerate output afterward.
- Existing user/unrelated changes in target file → preserve them and make the smallest contextual edit; do not reset/overwrite wholesale.
- Acceptance fails for a clearly introduced syntax/type error → repair directly, rerun, and keep scope bounded.
- Similar implementation hypothesis fails repeatedly → `fable-recover`; arbitrary "try again" loops are forbidden.
- Acceptance command passes but does not exercise the changed path → do not claim card completion; strengthen acceptance or hand off with that gap explicit.
## Invariants
- Card scope does not expand silently.
- User/unrelated changes are preserved.
- Source-of-truth is edited instead of derived output when applicable.
- Every meaningful mutation makes prior verification stale.
- No unrelated cleanup is bundled into a repair unless required for correctness.
- Acceptance evidence corresponds to the changed behavior.
## Failure Taxonomy
### Local implementation defect
Current edit causes a direct syntax/type/assertion failure. Fix within card.
### Hypothesis failure
Code is valid but behavior remains wrong. Revise hypothesis; repeated similar failures route to recovery.
### Hidden dependency
Correct behavior requires a component/contract outside the accepted card. Stop and replan scope.
### Harness/artifact mismatch
Commands execute stale build, wrong env, wrong branch, or generated artifact. Route to recovery when this obscures causal evidence.
### Acceptance weakness
Check passes without exercising the change. Strengthen proof before completion.
### Concurrent-work conflict
Target contains legitimate changes not owned by this card. Preserve them, narrow edit, or coordinate ownership; never overwrite for convenience.
## Anti-Patterns
- "while I'm here" refactors;
- editing generated files directly;
- broad search/architecture work after execution starts;
- retrying the same patch with superficial syntax changes;
- replacing an entire user-modified file for a small fix;
- accepting a build pass as functional proof when behavior is not exercised;
- adding abstraction/configurability not required by the card;
- hiding an architecture decision inside implementation.
## Execution Packet
```text
Card:
Owned scope:
Files/surfaces changed:
Behavior/contract delta:
Acceptance command/result:
Invariants checked:
Generated/config impacts:
Unexpected facts:
Residual risk:
Recommended verify scope:
```
## Completion Criteria
Execution completes when:
- the accepted behavior is implemented within explicit scope;
- immediate acceptance evidence is green or any proof gap is explicitly handed off;
- diff contains no accidental/unrelated mutation;
- new architecture decisions have not been smuggled into implementation;
- the verification specialist has a precise map of what changed and what remains to falsify.
## Progressive Resources
- Deep containment guide: `references/bounded-mutation-and-failure-classification.md`
- Existing containment reference: `references/mutation-containment.md`
- Example: `examples/bounded-card-execution.md`
Referenced files: 7
fable-handoff7.73 KB
---
name: fable-handoff
description: "Compact session decisions, durable evidence, open blockers, and exact next actions into structured continuation state for cross-session resumption. Use when pausing a coding session, transferring context to another agent, summarizing long-running work, or creating durable continuation checkpoints — even if the user does not explicitly say \"fable-handoff\" (e.g. \"save context for next session\", \"create a handoff\", \"pause work here\", \"summarize where we left off\"). Do NOT use as a substitute for verifying code changes (use fable-verify)."
version: 1.3.0
pack: delivery
inputs:
- current_state
requires:
- session_context
produces:
- handoff_evidence
- continuation_state
gates:
- next_action_explicit
fallback: get-fable
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-release
- fable-loop
continuations:
- get-fable
lateral_peers:
- fable-memory
recovery: fable-recover
---
# Fable Handoff
Leave enough verified context that another agent can resume the work without replaying the whole conversation or rediscovering the repository.
## Mission
A handoff is not a summary. It is a **resumption contract**.
The receiving agent should know what was requested, what was actually changed, which claims are proven, which evidence is stale or missing, what is currently blocked, and the first safe action to take next.
## Activate When
- pausing a long-running task;
- transferring work to another agent/human/session;
- context is becoming too large and needs durable compaction;
- a release/PR remains incomplete but execution must stop;
- an automation/cowork loop needs a durable checkpoint.
## Do Not Activate When
- there is nothing durable to preserve beyond a trivial completed turn;
- the next action itself is unclear because planning/diagnosis is unfinished;
- a handoff would be used to hide missing verification or unresolved failure.
## Continuity Classification
| State | Handoff emphasis |
| --- | --- |
| Clean completed card | exact evidence + next card |
| Dirty worktree | owned vs unrelated changes + safe resume point |
| Failed attempt | reproduction, attempts, revised/remaining hypotheses |
| Review pending | diff/base, findings, verification state |
| Release pending | candidate SHA/version, channel states, external blockers |
| Research/discovery pending | measured facts, unresolved load-bearing questions |
| Delegated work pending | worker ownership, returned/missing packets |
## Protocol
### Stage 1 — Reconstruct the durable task
Capture the original/current user outcome in one concise statement. Separate it from implementation tactics discovered later.
### Stage 2 — Snapshot exact repository/work state
Record when available:
- branch/worktree and candidate commit SHA;
- active card/phase;
- files/surfaces intentionally changed;
- unrelated user changes that must be preserved;
- generated artifacts/state files intentionally not committed;
- open PR/issues/release/tag identifiers.
### Stage 3 — Separate facts, decisions, and guesses
Use distinct buckets:
- **Measured facts**: observed repository/runtime/provider evidence;
- **Decisions**: selected architecture/contract and why;
- **Inferences**: plausible but not yet proven;
- **Rejected hypotheses**: avoid repeating failed reasoning;
- **Unresolved questions**: only those that can change the next action.
### Stage 4 — Record evidence with freshness
For each important claim record:
- command/probe/source;
- result;
- relevant mutation/commit/artifact;
- whether it is **fresh**, **stale**, or **not checked**.
Never turn `NOT_CHECKED` into prose that sounds like PASS.
### Stage 5 — Record failures and attempted repairs
Include only attempts that teach the receiver something. State what was tried, what happened, and why that changed or did not change the hypothesis.
### Stage 6 — State blockers precisely
A blocker must say what external fact/permission/environment is missing and what becomes possible when it is resolved.
Bad: `CI issue`.
Good: `GitHub Dependency Review fails because Dependency Graph is disabled for the repository; code/security scan itself did not report a dependency vulnerability.`
### Stage 7 — Give one next safe action
The first next action should be executable and justified by current evidence. If several actions are independent, list them after naming the primary one.
### Stage 8 — Sanitize
Do not copy secrets, tokens, private keys, raw credentials, unnecessary personal data, or giant logs. Refer to secure credential mechanisms rather than values.
## Decision Rules
- Prefer exact identifiers (branch, SHA, PR, file path, command) over narrative descriptions.
- Preserve negative knowledge: a disproved theory prevents repeated waste.
- Evidence freshness must survive compaction; old green tests cannot be summarized as current green.
- If the worktree is dirty, distinguish owned task changes from unrelated/user-owned changes.
- If next action depends on an unresolved decision, route to planning/recovery first rather than inventing a handoff action.
- Conversation history is optional context; the handoff must stand alone.
- Do not copy full logs when the error signature + command + key lines are sufficient.
## Invariants
- Handoff is truthful about what is and is not proven.
- Receiving agent can locate the exact work without guessing names/branches/files.
- Secrets are never persisted in the handoff.
- Unrelated user changes are explicitly protected.
- The next action cannot silently require context omitted from the handoff.
## Failure Taxonomy
### Narrative dump
Too much chat chronology, too little executable state. Rebuild around task/state/evidence/next action.
### False freshness
Old proof is summarized as current. Attach mutation/SHA/artifact context and downgrade stale evidence.
### Context starvation
Receiver must rediscover architecture or why a decision was made. Add the smallest load-bearing evidence/decision rationale.
### Secret leakage
Credential appears in summary/log. Remove value and reference secure source only.
### Ambiguous ownership
Dirty files are listed without saying which belong to task/user/other work. Clarify before handoff.
### Fake closure
Handoff says "done" while external gate/review/provider evidence is missing. Use explicit pending/blocker state.
## Anti-Patterns
- copying the entire conversation;
- dumping terminal output without interpretation;
- omitting failed attempts so next agent repeats them;
- saying "tests passed" without commit/mutation freshness;
- storing tokens because "the next agent will need them";
- vague next steps such as "continue fixing";
- losing user constraints/preferences that materially affect implementation;
- claiming release/publication from a prepared artifact only.
## Handoff Packet
```text
Goal:
Current branch/worktree/SHA:
Active phase/card:
Intentional changes:
Protected unrelated changes:
Measured facts:
Decisions + rationale:
Rejected hypotheses / failed attempts:
Evidence:
- claim → command/source → result → fresh/stale/not checked
Open findings/blockers:
External states (PR/CI/release/registry):
Primary next safe action:
Secondary independent actions:
Do not repeat / do not change:
```
## Completion Criteria
A handoff is complete when a zero-context receiving agent can:
- identify the exact task and repository state;
- distinguish facts/decisions/inferences;
- know which evidence is fresh;
- avoid repeating disproved attempts;
- preserve unrelated work and credentials safely;
- execute the next action without asking for missing basic context.
## Progressive Resources
- Deep guide: `references/resumability-and-context-compaction.md`
- Existing schema: `references/continuity-schema.md`
- Template: `templates/handoff-state.template.md`
- Example: `examples/cross-session-handoff.md`
Referenced files: 7
fable-loop8.2 KB
---
name: fable-loop
description: "Execute bounded recurring polling loops, CI build babysitting, interval-based status monitors, and self-paced test cycles with explicit timeouts and backoff. Use when monitoring async CI/CD pipelines, polling external service status, running interval test watches, or tracking long-running jobs — even if the user does not explicitly say \"fable-loop\" (e.g. \"babysit this build\", \"poll until completed\", \"watch CI status\", \"wait for deployment\"). Do NOT use for synchronous one-shot commands or unbounded infinite polling loops."
version: 1.3.0
pack: system
inputs:
- loop_condition
requires:
- exit_criteria
produces:
- loop_receipt
gates:
- budget_bounded
- exit_condition_explicit
fallback: fable-recover
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-run
continuations:
- fable-verify
- fable-handoff
lateral_peers:
- fable-run
recovery: fable-recover
---
# Fable Loop
Repeat a check only when repetition can reveal new state, with explicit termination, backoff, error classification, and zero ambiguity about why the loop stopped.
## Mission
Polling is not "run the same command until it turns green." A good loop models a changing external condition, distinguishes pending from failure, respects rate/cost budgets, and terminates on success, terminal failure, cancellation, or exhausted budget.
If repeated execution cannot produce new information, the task belongs in diagnosis—not a loop.
## Activate When
- waiting for CI/deployment/build/batch-job state to change;
- monitoring a bounded asynchronous operation;
- checking eventually-consistent external state;
- repeating a safe measurement while a known process progresses;
- running a finite stabilization sample where repetition itself is the measurement.
## Do Not Activate When
- one command/probe is enough (`fable-run`/`fable-verify`);
- the same deterministic failure is repeating with no external state change (`fable-recover`);
- user expects a future notification/scheduled task rather than an in-session bounded loop;
- polling would create repeated non-idempotent side effects;
- no termination criteria or budget can be defined.
## Loop Classification
| Loop type | Success/terminal semantics |
| --- | --- |
| CI/build | pending → success or terminal failure/cancelled |
| Deployment | progressing → healthy/rolled back/failed |
| Batch job | queued/running → completed/failed |
| Eventual consistency | old state → expected state within deadline |
| Rate-limited API | pending/retryable → success or terminal auth/schema error |
| Stabilization sampling | N bounded observations → distribution/variance verdict |
## Protocol
### Stage 1 — Define the state machine
Before iteration, enumerate:
- success state;
- pending/retryable states;
- terminal failure states;
- malformed/unknown states;
- cancellation condition.
A loop that treats every non-success as "try again" is unsafe.
### Stage 2 — Set budgets
Define at least:
- maximum elapsed time/deadline;
- maximum iterations or request budget where relevant;
- initial interval/backoff policy;
- maximum interval;
- API/cost/rate constraints.
Use the stricter bound when several apply.
### Stage 3 — Check idempotency and side effects
Polling operation should be read-only/idempotent. If the endpoint/command triggers work, separate trigger from status observation and ensure retries cannot duplicate the action.
### Stage 4 — Execute one iteration and classify result
Record:
- iteration/time;
- observed state/value;
- transport/command result;
- classification: success / pending / retryable error / terminal failure / unknown;
- next delay/reason.
### Stage 5 — Apply backoff intelligently
Use a fixed interval when the expected update cadence is known and inexpensive; exponential/backoff+jitter when rate limits/transient service errors matter.
Respect explicit `Retry-After`/provider guidance where applicable. Do not make sub-second aggressive calls to a slow external job just because tools allow it.
### Stage 6 — Stop early on terminal information
Exit immediately on:
- success;
- explicit failed/cancelled state;
- non-retryable auth/schema/permission error;
- user cancellation;
- budget/deadline exhaustion.
Do not consume remaining iterations after the outcome is already known.
### Stage 7 — Detect lack of progress
When a status includes progress/version/timestamp, compare across iterations. A long unchanged state near/after expected SLA may become a diagnostic signal rather than permission to extend budget automatically.
### Stage 8 — Produce an honest receipt
Final receipt includes stop reason, elapsed time, iterations, last state, transient errors, and whether the condition was actually satisfied.
Timeout is not success. "Still running" is not failure unless the contract/deadline says so.
## Decision Rules
- Deterministic repeated failure with no changing external state → stop loop and recover.
- Retryable transport error may continue within budget; auth/permission/schema errors usually terminate until configuration changes.
- Treat provider `Retry-After` or job-recommended poll interval as a lower bound where applicable.
- Never repeat a non-idempotent trigger as a status poll.
- If each iteration costs meaningful money/quota, include cost/request budget, not only time.
- If expected completion exceeds the current session/task model, create a scheduled/condition-watch mechanism when available rather than pretending an in-session loop can run indefinitely.
- A loop may end `INCOMPLETE/TIMEOUT`; do not extend bounds silently just to obtain green.
- If progress is unchanged and deadline still distant, continue according to policy without noisy user updates; surface only meaningful state changes for long interactive runs.
## Invariants
- Exit criteria and budgets are explicit before looping.
- Each iteration is safe/idempotent or side-effect semantics are explicitly controlled.
- Terminal failures stop immediately.
- Pending and failure are distinct states.
- Backoff/rate limits are respected.
- Receipt records actual stop reason; no timeout-to-pass conversion.
## Failure Taxonomy
### Infinite/implicit loop
No enforceable budget. Reject and define bounds.
### Retry-all-errors
Auth/schema/permission/terminal job failures are treated as transient. Classify and stop appropriately.
### Duplicate side effect
Polling call re-triggers work. Separate status endpoint/idempotency key or stop.
### Thundering poll
Interval too aggressive for service/job cadence. Back off/respect provider guidance.
### False timeout diagnosis
Job still pending within expected SLA but loop labels it broken. Report timeout/incomplete separately from product failure.
### Stuck progress
State never changes when it should. Hand to recovery/operations diagnosis rather than extending forever.
### Session mismatch
Requested monitoring lasts beyond available execution window. Use scheduling/condition-watch capability when supported.
## Anti-Patterns
- raw `while true`/unbounded sleep loops;
- retrying every error class;
- polling by repeatedly re-submitting the job;
- extending timeout until success after each miss;
- 1-second polling for a 10-minute deployment;
- declaring failure just because status is pending;
- declaring success because the loop ended cleanly;
- hiding transient errors from the final receipt;
- using loops to avoid diagnosing deterministic repeated failures.
## Loop Receipt
```text
Condition/state machine:
Success / terminal failure states:
Time + iteration/request budgets:
Interval/backoff policy:
Idempotency/side-effect check:
Iterations: time → observed state → classification
Transient errors:
Stop reason:
Condition satisfied? yes/no
Final state:
Next action if incomplete/failed:
```
## Completion Criteria
Loop completes when:
- state machine, success/failure semantics, and budgets were explicit;
- polling respected idempotency/rate/cost constraints;
- terminal information stopped the loop early;
- timeout/stuck state is represented honestly;
- final receipt shows exactly why looping ended and what should happen next.
## Progressive Resources
- Deep guide: `references/polling-state-machines-and-backoff.md`
- Existing guidelines: `references/loop-control-guidelines.md`
- Example: `examples/poll-build-job.md`
Referenced files: 7
fable-memory7.25 KB
---
name: fable-memory
description: "Manage persistent file-based memory, indexing cross-session user preferences, feedback, and architectural constraints in structured MEMORY.md stores. Use when recording user feedback, storing project conventions, recalling cross-session architectural constraints, or indexing durable project memory — even if the user does not explicitly say \"fable-memory\" (e.g. \"remember this preference\", \"save to memory\", \"what are my project rules\", \"recall user feedback\"). Do NOT use for volatile task state stored in .fable/state.json."
version: 1.3.0
pack: system
inputs:
- memory_fact
requires:
- fact_metadata
produces:
- memory_record
- updated_index
gates:
- single_fact_file
- index_synced
fallback: fable-discover
mutatesWorkspace: true
parallelSafe: true
neural_links:
precursors:
- fable-discover
continuations:
- fable-plan
- fable-discover
lateral_peers:
- fable-handoff
recovery: fable-recover
---
# Fable Memory
Persist only facts that deserve to survive the session, with provenance and contradiction handling strong enough that old memory does not silently become a new source of error.
## Mission
Memory is not a chat archive. It should reduce rediscovery without freezing temporary guesses into permanent instructions.
A durable record needs scope, provenance, confidence, freshness, and a way to be superseded when the user or repository changes.
## Activate When
- the user explicitly asks to remember a durable preference/constraint;
- a stable project convention or architectural decision will materially affect future work;
- cross-session continuity repeatedly depends on the same fact;
- a prior durable fact needs correction, supersession, or retrieval.
## Do Not Activate When
- the information is temporary execution state (`.fable/state`, ledger, handoff);
- it is a one-off conversational detail unlikely to affect future decisions;
- it is a guess/inference not yet stable enough to persist;
- it contains a password, token, private key, session cookie, or other secret.
## Memory Classification
| Type | Example | Durability rule |
| --- | --- | --- |
| User preference | preferred package manager | persist when explicit/stable |
| Project constraint | runtime Bun >=1.3 | bind to project/source |
| Architecture decision | use outbox for events | store rationale + supersession path |
| Repeated workflow rule | verify package clean-install before release | persist if project-specific and durable |
| External fact | API behavior/version | usually cite/research at use time; avoid timeless storage |
| Temporary state | current failing test | handoff/state, not memory |
| Sensitive credential | API token | never memory |
## Protocol
### Stage 1 — Decide whether it deserves memory
Ask:
- will this likely matter in another session?
- is it stable, explicit, or source-backed?
- is there a narrower existing record to update?
- could persistence create privacy/security risk?
If not durable, leave it in current context/handoff only.
### Stage 2 — Normalize the fact
Store one coherent decision/fact with:
- canonical statement;
- scope (user/project/repo/subsystem);
- type;
- source/provenance;
- confidence/status;
- created/updated timestamp if supported;
- supersedes/superseded-by relation when relevant.
### Stage 3 — Check conflicts before write
Search for existing records with same subject. Compare:
- exact agreement → update provenance/freshness rather than duplicate;
- narrower/wider scope → preserve both only if scopes are genuinely different;
- contradiction → do not keep two active truths; resolve or mark conflict explicitly.
### Stage 4 — Write atomically and index
Update one canonical fact record and synchronize index/catalog. Preserve unrelated records.
### Stage 5 — Validate retrieval meaning
Read back the record as a future agent would. Ensure it does not overgeneralize:
- `use Bun in get-fable` is not `user always uses Bun everywhere`;
- `API v4 did X in Aug 2026` is not an eternal API guarantee.
### Stage 6 — Apply memory skeptically on retrieval
When a remembered fact affects current work:
- confirm scope matches;
- prefer current user/repository evidence over old memory;
- revalidate time-sensitive external facts;
- treat explicit new instruction as superseding old preference where applicable.
## Decision Rules
- User's current explicit statement outranks stored preference.
- Repository/config/source evidence outranks contradictory memory about repository state.
- Time-sensitive external facts should be researched again rather than trusted indefinitely.
- A memory can be historical without remaining active; use superseded status instead of deletion when history matters.
- Do not infer broad personal preferences from a single project choice.
- Never store credentials even when the user asks to "remember the token"; reference secure credential management instead.
- One fact per record is a maintainability rule, but related rationale/source may accompany the fact.
## Invariants
- No secrets enter durable memory.
- Active memory contains no unresolved contradictory truths without explicit conflict status.
- Scope is explicit enough to prevent accidental generalization.
- Provenance is retained for consequential constraints/decisions.
- Current evidence/instruction can supersede memory.
- Index and underlying record remain consistent.
## Failure Taxonomy
### Duplicate memory
Same fact exists twice with minor wording differences. Merge/update canonical record.
### Scope leak
Project-specific choice is treated as global user preference. Narrow the scope.
### Stale memory
Repository/user/external reality changed. Supersede or revalidate before use.
### Contradictory active facts
Two records give incompatible active guidance. Resolve with current source/user evidence.
### Inference hardened into fact
Agent stored an assumption as durable truth. Downgrade/remove and require evidence.
### Secret persistence
Sensitive value appears in memory. Remove exposure and handle credential rotation/security as appropriate.
## Anti-Patterns
- saving every conversation detail "just in case";
- storing temporary errors/active todos as durable preference;
- broadening `in this repo` into `always`;
- persisting current API/version claims with no expiry/provenance;
- creating a new fact file instead of updating a conflicting one;
- storing a token because future automation may need it;
- treating memory as more authoritative than current repository evidence;
- silently deleting historical decisions when supersession context matters.
## Memory Record
```text
Subject/fact:
Scope:
Type:
Status: active | superseded | conflicted
Source/provenance:
Confidence:
Rationale/context:
Supersedes / superseded by:
Freshness/revalidation note:
```
## Completion Criteria
Memory work completes when:
- fact is genuinely durable and safely scoped;
- duplicates/conflicts were checked;
- record has enough provenance to evaluate later;
- index is synchronized;
- sensitive/temporary data was excluded;
- a future agent can retrieve and apply the fact without overgeneralizing it.
## Progressive Resources
- Deep guide: `references/durable-fact-lifecycle.md`
- Existing protocol: `references/memory-management-protocol.md`
- Template: `templates/memory-fact.template.md`
- Example: `examples/recording-user-preference.md`
Referenced files: 7
fable-plan8.22 KB
---
name: fable-plan
description: "Convert discovery evidence into bounded, testable work cards with explicit acceptance criteria and architectural invariants. Use when designing multi-file features, planning complex refactors, decomposing architectural migrations, or structuring multi-step implementation tasks — even if the user does not explicitly say \"fable-plan\" (e.g. \"plan this feature\", \"design the architecture\", \"break this task down\", \"create an implementation plan\"). Do NOT use when load-bearing facts remain unknown (use fable-discover first) or for trivial single-scope bug fixes (use fable-tdd or fable-execute)."
version: 1.3.0
pack: core
inputs:
- requirements
- evidence
requires:
- load_bearing_unknowns_resolved
produces:
- accepted_card
- acceptance_condition
gates:
- bounded_scope
- named_acceptance
fallback: fable-discover
mutatesWorkspace: false
parallelSafe: false
neural_links:
precursors:
- fable-discover
- fable-research
continuations:
- fable-tdd
- fable-delegate
- fable-execute
lateral_peers:
- fable-config
recovery: fable-recover
---
# Fable Plan
Turn grounded evidence into an executable change strategy with explicit dependencies, risk boundaries, acceptance proof, and integration ownership.
## Mission
A plan is not a list of files to edit. It is a set of decisions that makes implementation boring enough to execute safely.
The planner should remove architectural uncertainty, expose order constraints, define what "done" means, and separate work only where the pieces can genuinely be reasoned about and verified independently.
## Activate When
- the task spans multiple files/components or introduces a new contract;
- a migration, schema change, public API change, or cross-package refactor is required;
- implementation requires sequencing or rollback reasoning;
- parallel workers may help but ownership/integration must be designed;
- requirements are understood but not yet decomposed into verifiable work.
## Do Not Activate When
- load-bearing facts are still unknown (`fable-discover`/`fable-research`);
- the change is truly bounded to one obvious edit with clear acceptance (`fable-execute`);
- the immediate job is to reproduce a behavior in a test (`fable-tdd`).
## Change Classification
Classify before decomposing.
| Change type | Main planning concern |
| --- | --- |
| Bounded behavior | acceptance and regression surface |
| Cross-module feature | contracts + integration points |
| Refactor | invariant preservation + staged compatibility |
| Data/schema migration | forward/backward compatibility + rollback |
| Public API/CLI | compatibility + versioning + docs/tests |
| Dependency upgrade | behavior delta + transitive risk |
| Security-sensitive | trust boundaries + fail-closed behavior |
| Release/distribution | artifact boundaries + external verification |
## Planning Protocol
### Stage 1 — Restate the contract
Capture:
- requested outcome;
- non-goals;
- evidence already established;
- externally visible behavior that must not regress;
- unresolved items and why they are non-blocking.
If a blocking unknown remains, stop and route back to discovery/research.
### Stage 2 — Map the change graph
Identify:
- components touched;
- contracts between them;
- data/state transitions;
- migration or compatibility edges;
- shared integration hotspots;
- generated artifacts and their sources;
- verification surfaces.
Represent ordering as dependencies, not just numbered tasks.
### Stage 3 — Choose seams for decomposition
A card should have one coherent reason to change and one clear acceptance story.
Do **not** require disjoint files as a universal rule. Two cards may touch one integration file sequentially if ownership/order is explicit. Conversely, two cards touching different files may still be unsafe to parallelize if they change one shared invariant or contract.
### Stage 4 — Plan risk-first
Move high-uncertainty or irreversible decisions earlier where possible:
- prove migration shape before broad data rewrite;
- prove compatibility adapter before deleting legacy path;
- prove third-party API contract before wiring all call sites;
- introduce observability before risky runtime changes.
### Stage 5 — Define each card
Each card must state:
- objective and user-visible effect;
- owned scope and prohibited scope;
- dependencies/preconditions;
- implementation notes only where they constrain correctness;
- acceptance evidence/commands;
- rollback or recovery note when failure is costly;
- handoff/integration expectation.
### Stage 6 — Define integration ownership
Name who/what proves the combined result. Worker-local green is not enough for cross-card behavior.
### Stage 7 — Check the plan against failure
Ask:
- What if card 2 fails after card 1 lands?
- Can old and new formats coexist during migration?
- Is any step irreversible?
- Does rollback preserve data/contracts?
- Are tests capable of detecting the main regression?
- Does parallel execution create merge or semantic conflicts?
## Decision Rules
- Unknown architecture → `fable-discover`; unknown external contract → `fable-research`.
- Behavior change with a testable contract → route the first implementation card through `fable-tdd`.
- Parallelize only when dependencies, ownership, shared invariants, and integration points make concurrency safe.
- Prefer staged compatibility over flag-day replacement when old/new callers or persisted data may coexist.
- If a card cannot name falsifiable acceptance evidence, it is not ready.
- If planning produces more than roughly 5 tightly interdependent cards, consider milestone boundaries or a different architecture rather than mechanically splitting more.
- Generated files belong to the card that changes their source/generator, not as independent manual edits.
## Invariants
- Every card traces back to a requirement or necessary enabling constraint.
- No blocking architectural unknown is hidden inside an execution card.
- Acceptance checks prove behavior, not merely file existence.
- Integration verification has an owner.
- Destructive/irreversible steps have explicit safeguards or rollback reasoning.
- Parallelism is justified by dependency analysis, not by file count alone.
## Failure Taxonomy
### Hidden dependency
A card unexpectedly requires another component. Update the dependency graph; do not let execution silently broaden scope.
### False independence
Cards touch disjoint files but share one invariant/contract. Serialize or define a compatibility seam.
### Acceptance weakness
A proposed check can pass while behavior is still broken. Strengthen the acceptance condition before execution.
### Migration hazard
Old/new states can coexist in production but plan assumes atomic cutover. Add compatibility, backfill, or staged rollout.
### Scope explosion
A card keeps discovering architectural work. Stop, return to discovery/plan, and redraw boundaries.
## Anti-Patterns
- one card per file;
- requiring disjoint files even when sequential integration is safer;
- parallelizing based only on "different files";
- acceptance such as "tests pass" without naming relevant behavior;
- hiding migrations/rollbacks inside implementation prose;
- treating documentation/generated output as separate cards when they are part of one behavior change;
- producing a long task list with no dependency model;
- over-specifying code line-by-line before execution evidence exists.
## Work Card Template
```text
Card: <outcome>
Why: <requirement/risk>
Depends on:
Owns:
Must not change:
Behavior/invariant:
Implementation constraints:
Acceptance evidence:
Rollback/recovery:
Integration handoff:
Parallel-safe with:
```
## Completion Criteria
Planning is complete when:
- blocking unknowns are gone;
- dependency/order constraints are explicit;
- cards are coherent and bounded by behavior, not arbitrary file slicing;
- each card has falsifiable acceptance evidence;
- migration/compatibility/rollback concerns are addressed where relevant;
- parallel work, if proposed, has explicit semantic ownership and integration proof.
## Progressive Resources
- Deep decomposition guide: `references/risk-and-dependency-decomposition.md`
- Existing decomposition rules: `references/decomposition-rules.md`
- Template: `templates/work-card.template.md`
- Example: `examples/multi-file-migration-plan.md`
Referenced files: 7
fable-recover8.96 KB
---
name: fable-recover
description: "Diagnose repeated command failures, stale build caches, branch drift, or contradictory evidence before attempting further code edits. Use when commands fail repeatedly, tests stay red after attempted fixes, build output contradicts source code, or the execution path is confused — even if the user does not explicitly say \"fable-recover\" (e.g. \"still failing\", \"why did this fail again\", \"stuck in a failure loop\", \"diagnose this error\"). Do NOT use for routine first-time test failures in fresh TDD (use fable-tdd)."
version: 1.3.0
pack: core
inputs:
- failure_evidence
requires:
- failure_streak
produces:
- revised_hypothesis
- bounded_repair
gates:
- diagnosis_changed
fallback: fable-discover
mutatesWorkspace: false
parallelSafe: false
neural_links:
precursors:
- fable-tdd
- fable-execute
- fable-verify
continuations:
- fable-discover
- fable-plan
- fable-execute
lateral_peers:
- fable-discover
recovery: fable-discover
---
# Fable Recover
Stop spending mutations on a failing hypothesis. Rebuild causal confidence before changing code again.
## Mission
Recovery exists for the moment an agent is most likely to become expensive and irrational: the same task has failed more than once, output contradicts expectations, or edits appear to have no effect.
The goal is not to "try something different." The goal is to explain why previous attempts failed, falsify competing causes, and issue one repair that is justified by new evidence.
## Activate When
- the same command/test/behavior fails after two materially similar implementation attempts;
- output appears stale or unaffected by known source changes;
- tests/runtime disagree;
- evidence contradicts a load-bearing assumption;
- failures alternate or depend on timing/environment;
- the agent is about to repeat a command/patch without new diagnostic information.
## Do Not Activate When
- first failure is a trivial syntax/type error introduced by the current edit;
- architecture is simply unknown and no repeated failure occurred (`fable-discover`);
- expected RED in TDD is being observed correctly;
- a narrow review finding already has an obvious bounded repair.
## Failure Classification
Start by classifying the evidence, not the code.
| Class | Typical signals | First probes |
| --- | --- | --- |
| Harness | assertion never reached, fixture/mock/setup error | prove test path and fixture |
| Environment | local/CI/OS/env-specific | versions, env, cwd, permissions |
| Artifact/cache | source changes not reflected | entrypoint, build timestamps, cache, dist |
| Execution path | edited code never executes | tracing/instrumentation/registration |
| Dependency/version | unexpected API/runtime semantics | lockfile, resolved version, official source |
| Data/state | only some fixtures/accounts/orders fail | minimal failing state, persistence boundaries |
| Concurrency/timing | intermittent/order-sensitive | deterministic coordination, shared state |
| Product logic | harness/path proven, assertion consistently wrong | isolate algorithm/branch |
| Invariant/design | local fixes move failure elsewhere | identify violated cross-component rule |
Do not jump to product logic until cheaper external explanations are falsified.
## Recovery Protocol
### Stage 1 — Freeze mutation
No new production edits until the diagnosis changes. Preserve the failing state and collect exact evidence.
Record:
- command/action;
- exact error/output;
- workspace/commit/mutation generation;
- attempts already made and what differed;
- expected observation.
### Stage 2 — Reproduce minimally
Find the smallest reliable reproduction. If the failure is flaky, capture seeds/order/time/environment and work on determinism before another fix.
### Stage 3 — Build a hypothesis queue
Create 2-5 plausible causes ranked by:
- ability to explain all observed evidence;
- probability given recent changes;
- cost/safety of falsification.
Each hypothesis must predict an observation that would distinguish it.
Bad: "maybe cache."
Good: "CLI executes stale `dist/cli.js`; if true, source timestamp will be newer than dist and direct source invocation will show new behavior."
### Stage 4 — Walk the attribution ladder
Use the cheapest separating probes first:
1. **Harness** — does the test/probe reach the intended assertion/path with realistic inputs?
2. **Environment/artifact** — correct branch, cwd, env, version, build, cache, process?
3. **Execution path** — is edited code actually reached? which implementation is registered?
4. **Data/dependency** — does input/version/state differ from assumptions?
5. **Product logic** — with above proven, isolate the wrong branch/algorithm.
6. **Invariant/design** — if local logic is individually reasonable but system remains wrong, identify the violated system contract.
The ladder is guidance, not ritual. Skip a rung only when existing evidence already proves it.
### Stage 5 — Instrument or bisect when observation is weak
Use narrow temporary diagnostics, binary search/bisect, toggling one variable, or comparing known-good/bad states.
Change one diagnostic dimension at a time so the result is interpretable.
### Stage 6 — Falsify, do not accumulate guesses
After each probe:
- reject hypothesis;
- strengthen hypothesis;
- or revise the queue.
Do not keep contradicted explanations alive as "maybe still related."
### Stage 7 — Form the revised diagnosis
A valid diagnosis explains:
- why the observed failure occurred;
- why previous attempts did not fix it;
- what evidence distinguishes it from alternatives;
- what smallest repair should change the outcome.
### Stage 8 — Issue one bounded repair
Return to `fable-execute` or `fable-tdd` with one repair and one expected proof. If diagnosis reveals architecture uncertainty, route to plan/discover instead.
## Decision Rules
- Never repeat an unchanged failed command unless a named environmental/state variable changed or the rerun is explicitly measuring nondeterminism.
- Do not delete caches/build artifacts reflexively before recording evidence; destructive cleanup can erase the clue that proves staleness.
- If source changes have no runtime effect, prove the executed artifact/path before editing logic again.
- If CI-only failure exists, compare environment/version/parallelism first; do not assume CI is "random."
- If failure is data-specific, minimize the failing data/state before broad refactor.
- If a dependency/version hypothesis emerges, route external semantic verification to `fable-research`.
- If local patches shift failure between components, suspect a shared invariant/design and return to `fable-plan`.
- Similar failed fixes count as a failure streak; superficial syntax changes do not reset diagnostic responsibility.
## Invariants
- Recovery is diagnostic/read-only until a revised diagnosis exists.
- Every new probe is chosen to distinguish hypotheses.
- Contradicted hypotheses are removed.
- Diagnostic instrumentation is temporary and cleaned after repair.
- Final repair is bounded and tied to a predicted observable outcome.
## Failure Taxonomy of Recovery Itself
### Blind cleanup
Cache/build reset makes problem disappear but root cause is unknown. Record as unresolved unless causal evidence is obtained.
### Hypothesis sprawl
Long list of possibilities with no discriminating probes. Rank and test the cheapest separator.
### Mutation during diagnosis
Agent edits product while still uncertain, invalidating the failing state. Revert/restore diagnostic baseline where safe and restart evidence collection.
### Confirmation bias
Only probes supporting first theory are run. Add at least one falsifier for the leading hypothesis.
### Non-minimal reproduction
Huge suite/system creates too many confounders. Isolate smaller path before interpreting results.
## Anti-Patterns
- third/fourth patch with same causal theory;
- "clear cache and see" without recording before/after evidence;
- rerunning flaky command until green;
- blaming environment without comparing environments;
- adding broad logging everywhere;
- changing multiple diagnostic variables at once;
- preserving disproved assumptions;
- solving symptom while unable to explain previous failures.
## Recovery Packet
```text
Failure/reproduction:
Attempts already made:
Hypothesis queue:
Probe → observation → hypothesis effect:
Revised diagnosis:
Why prior attempts failed:
Bounded repair:
Expected proof after repair:
Residual uncertainty:
Next Skill:
```
## Completion Criteria
Recovery completes only when:
- a reliable enough reproduction exists or nondeterminism is explicitly characterized;
- leading alternative causes were falsified with evidence;
- diagnosis materially differs from the failed assumption/attempt;
- proposed repair is bounded and predicts a concrete changed observation;
- execution can resume without another blind retry.
## Progressive Resources
- Deep guide: `references/diagnostic-falsification-playbook.md`
- Existing ladder: `references/attribution-ladder.md`
- Example: `examples/recovering-stale-test-cache.md`
Referenced files: 7
fable-release9.48 KB
---
name: fable-release
description: "Audit and certify repository merge and release readiness against required quality gates, clean git working trees, and verified distribution artifacts. Use when preparing a release tag, validating release checklist criteria, publishing npm/PyPI packages, or certifying a branch for merge — even if the user does not explicitly say \"fable-release\" (e.g. \"prepare the release\", \"is this ready to merge\", \"ship to production\", \"publish the package\"). Do NOT use when verification is stale, failing, or missing (use fable-verify first)."
version: 1.3.0
pack: delivery
inputs:
- completion_evidence
requires:
- clean_worktree
produces:
- release_readiness
gates:
- required_checks_pass
- no_blocking_findings
fallback: fable-verify
mutatesWorkspace: false
parallelSafe: false
neural_links:
precursors:
- fable-verify
- fable-review
- fable-security
continuations:
- fable-handoff
lateral_peers:
- fable-handoff
recovery: fable-recover
---
# Fable Release
Prove that the exact commit and artifact users will receive are ready to ship, then distinguish readiness from actual distribution.
## Mission
A release is an external state transition. Source tests can be perfect while the package omits files, the tag points at another commit, the registry still serves an old version, or a published artifact cannot start in a clean environment.
This Skill closes that gap by tying release claims to exact version, commit, artifact, workflow, tag, and registry evidence.
## Activate When
- preparing a branch for merge or a version for release;
- creating/pushing a version tag;
- publishing npm/package/plugin artifacts;
- finalizing a GitHub Release;
- verifying that a distribution channel actually serves the intended version;
- assessing whether release evidence is current after last-minute changes.
## Do Not Activate When
- implementation is still changing (`fable-execute`/`fable-tdd`);
- functional verification is incomplete (`fable-verify`);
- blocking review/security findings remain;
- the user has not authorized an irreversible publish action. Readiness checks may run; publishing itself requires the applicable authorization.
## Release Classification
| Release type | Extra concerns |
| --- | --- |
| Merge only | branch/base freshness, required CI, review |
| Package registry | package contents, clean install, registry verification |
| Git tag/GitHub Release | tag→SHA binding, notes/assets, draft/prerelease state |
| Plugin/marketplace | manifest/version parity, submission/approval state |
| Migration-bearing release | rollout order, backward compatibility, rollback |
| Security-sensitive release | advisory/secrets/dependency gates, disclosure constraints |
## Protocol
### Stage 1 — Freeze the candidate
Identify the exact candidate:
- version;
- commit SHA;
- target branch/tag;
- distribution channels;
- required CI/review/security evidence.
Any source/config/package mutation after this point invalidates relevant release evidence and creates a new candidate.
### Stage 2 — Validate version semantics
Check:
- version does not already exist in target registry/tag namespace;
- SemVer matches actual compatibility impact;
- all canonical version locations/manifests agree;
- changelog/release notes describe user-visible changes accurately;
- prerelease state is intentional.
Do not choose patch/minor/major purely for convenience.
### Stage 3 — Reconfirm fresh quality gates
Verify required functional/build/review/security evidence belongs to the candidate SHA/mutation generation. If evidence was produced before candidate changes, rerun it.
### Stage 4 — Inspect the artifact boundary
Build/dry-run the exact artifact and inspect its manifest/content.
Check:
- required entrypoints/assets/manifests are present and non-empty;
- secrets, temp files, tests/fixtures/internal holdouts are excluded unless intentionally public;
- generated files are current;
- executable permissions/exports/bin paths are correct;
- dependency/runtime constraints are accurate.
### Stage 5 — Clean-environment smoke
Install/use the produced artifact outside the source checkout where feasible.
Prove at least:
- install succeeds;
- primary executable/import resolves;
- version/help/basic smoke works;
- documented quick-start path is not accidentally relying on repo-local files.
### Stage 6 — Verify repository release state
Before publishing:
- candidate commit is pushed;
- required CI on the candidate is green;
- tag does not exist or already points to exactly the intended SHA;
- release notes/assets target the same tag/SHA;
- branch is not known-broken relative to base.
### Stage 7 — Publish only through an authorized secure path
Prefer configured trusted publishing/OIDC/host automation over introducing long-lived tokens.
Do not print/store credentials or weaken security to make a release pass.
If authorization or authentication is absent, stop at **READY_NOT_PUBLISHED** with exact remaining action.
### Stage 8 — Verify distribution independently
After publish, query the external destination rather than trusting the publish command.
Examples of proof:
- registry reports expected version/dist metadata;
- clean install from registry succeeds;
- Git tag resolves to candidate SHA;
- GitHub Release is actually public, not draft;
- marketplace status reflects submitted/published state.
### Stage 9 — Post-release smoke and handoff
Run the user-facing install/invocation path from public distribution. Record rollback/deprecation/follow-up issues if observed.
## Release Verdicts
- **NOT_READY**: required deterministic/review/security/artifact gate fails.
- **READY_NOT_PUBLISHED**: candidate is sound but publish is unauthorized/unavailable/not requested.
- **PUBLISHED_UNVERIFIED**: publish command/workflow claims success but external distribution has not been independently confirmed.
- **RELEASED**: exact candidate is public through intended channels and independently verified.
Do not collapse these states into "done."
## Decision Rules
- A dirty worktree does not automatically fail if dirt is explicitly unrelated and release artifact is from a clean candidate SHA; however uncommitted release changes do fail readiness.
- Tag exists at different SHA → hard stop; never move/force a release tag casually.
- Package version already exists in immutable registry → choose a new valid version, do not overwrite.
- Dry-run contents differ from intended public surface → fix package configuration before tagging/publishing.
- Local global install works from repo link but clean tarball/registry install fails → NOT_READY.
- CI green on older commit → stale release evidence.
- GitHub Release draft exists → not public release evidence.
- Publish workflow green but registry not updated yet → PUBLISHED_UNVERIFIED until independently confirmed or known propagation policy is resolved.
- Marketplace requiring manual approval → report submission state accurately; never claim publication.
## Invariants
- Release version, commit, tag, artifact, and public metadata refer to one candidate.
- No credential is committed or logged as release evidence.
- Irreversible external writes require applicable authorization.
- Public release claims are externally verified.
- Last-minute mutation invalidates stale candidate evidence.
- Release notes do not claim maturity/features unsupported by fresh proof.
## Failure Taxonomy
### Artifact omission/pollution
Package misses runtime file or includes secrets/internal material. Fix boundary, rebuild, resmoke.
### Version drift
Manifests/changelog/CLI/tag disagree. Reconcile before release.
### Candidate drift
CI/evidence points to older SHA after release commit changed. Rerun gates.
### Tag mismatch
Existing tag points elsewhere. Stop; investigate rather than force-move.
### Publish/auth failure
Candidate may remain READY_NOT_PUBLISHED. Diagnose secure auth/workflow without embedding credentials.
### Registry/release mismatch
Publish reports success but public channel serves wrong/old artifact. Keep PUBLISHED_UNVERIFIED and investigate.
### Clean-install failure
Artifact depends on repo-local files/dev state. NOT_READY regardless of source tests.
## Anti-Patterns
- "tests pass, publish";
- tagging before inspecting artifact contents;
- using source checkout as the only install test;
- force-moving tags;
- adding tokens to config/docs to bypass trusted publishing;
- equating draft release with public release;
- trusting workflow success without registry/release lookup;
- updating docs to claim availability before public verification;
- claiming 100% maturity while changed behavioral evidence is stale.
## Release Attestation
```text
Version:
Candidate SHA:
Target channels:
Fresh quality gates:
Artifact manifest check:
Clean-install smoke:
Tag → SHA:
GitHub Release state:
Registry/marketplace state:
External verification:
Verdict: NOT_READY | READY_NOT_PUBLISHED | PUBLISHED_UNVERIFIED | RELEASED
Rollback/follow-up:
```
## Completion Criteria
The Skill's work completes when the requested release stage is accurately attested:
- readiness claims are tied to exact candidate evidence;
- artifact boundary and clean install are proven where applicable;
- tag/release/registry states are consistent;
- public distribution is independently verified before `RELEASED`;
- missing authorization/external gates are represented as explicit state, never guessed away.
## Progressive Resources
- Deep guide: `references/artifact-and-distribution-verification.md`
- Existing release gates: `references/release-gates.md`
- Example: `examples/release-readiness-audit.md`
Referenced files: 7
fable-research7.55 KB
---
name: fable-research
description: "Resolve current external facts, official documentation, library behaviors, and API contracts against primary sources before implementation. Use when consulting library documentation, checking breaking changes, investigating external APIs, or verifying framework versions — even if the user does not explicitly say \"fable-research\" (e.g. \"check the latest docs\", \"what is the API for X in version Y\", \"look up SDK specs\", \"verify library support\"). Do NOT use for repository-local code exploration (use fable-discover) or for writing implementation code directly (use fable-tdd or fable-execute)."
version: 1.3.0
pack: intelligence
inputs:
- research_query
requires:
- primary_sources
produces:
- research_evidence
- source_backed_facts
gates:
- primary_source_grounding
fallback: fable-discover
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-discover
continuations:
- fable-plan
- fable-tdd
- fable-execute
lateral_peers:
- fable-discover
recovery: fable-recover
---
# Fable Research
Resolve external facts that can change the implementation, using current primary evidence instead of model memory.
## Mission
Research is not link collection. The output must be a decision-ready statement tied to the exact version, environment, or contract the repository will use.
## Activate When
- SDK/API signatures, limits, defaults, lifecycle semantics, or compatibility may have changed;
- a design depends on current cloud/provider behavior;
- the repository pins a version that may differ from current docs;
- multiple official sources appear to disagree;
- an error may come from a documented breaking change or deprecation;
- a standard/RFC/security requirement must be interpreted precisely.
## Do Not Activate When
- the fact is repository-local (`fable-discover`);
- the implementation is already grounded and only needs execution (`fable-execute`/`fable-tdd`);
- the task is general brainstorming where no external claim changes the decision.
## Research Classification
Classify the question first.
| Question type | Best primary evidence | Extra risk |
| --- | --- | --- |
| API signature | versioned official docs + source/types | docs may show latest, repo pins older |
| Runtime behavior | official docs + upstream implementation/tests | marketing docs may omit edge behavior |
| Compatibility | release notes/changelog + version matrix | transitive dependency constraints |
| Standard/protocol | normative spec/RFC | examples may be non-normative |
| Cloud/product limit | current vendor docs | region/tier/account differences |
| Security guidance | official security docs/advisories | stale blog summaries |
| Deprecation/migration | migration guide + release notes | old and new APIs coexist |
## Protocol
### Stage 1 — Turn the task into answerable claims
Break a broad request into the smallest external claims that can affect design.
Bad: "Research the new SDK."
Good:
- Does version 4.2 expose streaming tool-call deltas?
- Which parameter enables them?
- Is the callback ordered?
- What minimum runtime version is required?
### Stage 2 — Bind to repository reality
Before accepting current docs as applicable, record:
- package/version actually used;
- runtime/language version;
- relevant feature flags/tier/region if applicable;
- whether the repository uses generated types or a wrapper that changes the public contract.
### Stage 3 — Use a source hierarchy
Prefer, in order when available:
1. normative specification or official versioned reference;
2. official upstream source/types/tests;
3. official release notes/migration guides/advisories;
4. vendor examples authored for the relevant version;
5. secondary sources only as leads.
Never use an unsourced search snippet as final evidence.
### Stage 4 — Reconcile version and source conflicts
If latest docs disagree with the pinned package:
- inspect versioned docs/release notes/source for the pinned version;
- state the delta explicitly;
- do not silently recommend latest syntax to an older lockfile.
If two official sources disagree, prefer the one closest to executable truth for the exact version, and report the conflict.
### Stage 5 — Separate fact from interpretation
Record:
- **Fact**: what the source establishes;
- **Applicability**: why it applies to this repo/version;
- **Interpretation**: what it means for the design;
- **Confidence**: measured / strongly supported / unresolved.
### Stage 6 — Stop when the decision is grounded
Do not continue reading once all load-bearing external claims are resolved and the next action is safe.
## Decision Rules
- If a fact could have changed since model training, verify it rather than recall it.
- If docs are unversioned and the repo pins an older version, inspect source/types/changelog for that exact version.
- If an official quickstart conflicts with a normative reference, do not flatten the disagreement; determine which governs the target behavior.
- If behavior depends on account/tier/region, label that dependency and avoid universal claims.
- If no primary source is available, report the evidence gap and use upstream code/types/tests as the next-best source; never invent missing parameters.
- If the answer changes architecture, hand off to `fable-plan`; if it simply confirms a bounded implementation contract, hand off to `fable-tdd` or `fable-execute`.
## Invariants
- Every load-bearing external claim has a source.
- Source applicability includes version/context, not just URL authority.
- Quotes/signatures are kept short and exact; conclusions are written in the agent's own words.
- Secondary sources do not override accessible primary sources.
- Conflicting evidence remains visible until resolved.
## Failure Taxonomy
### Freshness failure
The source is official but stale/deprecated. Find versioned/current material and release history.
### Version mismatch
The repo and docs describe different versions. Reconcile against the lockfile/package metadata.
### Authority mismatch
A blog/example contradicts normative docs or upstream source. Demote the weaker source.
### Context mismatch
The claim differs by region, tier, runtime, platform, or feature flag. Scope the conclusion.
### Interpretation ambiguity
The source is clear but its implication for the repository is not. Hand the unresolved design question to `fable-plan` rather than pretending the research answered it.
## Anti-Patterns
- asking a search engine for a signature and copying the snippet;
- citing the latest docs without checking the pinned version;
- collecting many links without a decision-ready conclusion;
- using model memory because the API "probably hasn't changed";
- quoting a source without saying why it applies;
- hiding official-source disagreement;
- continuing research after every load-bearing claim is resolved.
## Research Packet / Handoff
```text
Question:
Repository version/context:
Source-backed facts:
- fact → primary source → applicability
Conflicts/version deltas:
Interpretation for implementation:
Confidence / unresolved:
Recommended next Skill:
```
## Completion Criteria
Research is complete when:
- the exact external claim is answered for the repository's actual context;
- evidence is primary or the absence of primary evidence is explicit;
- version conflicts are reconciled;
- the implementation/design implication is stated without overclaiming;
- no load-bearing parameter or semantic remains ambiguous.
## Progressive Resources
- Deep guide: `references/source-reconciliation-playbook.md`
- Existing hierarchy: `references/primary-source-hierarchy.md`
- Example: `examples/research-api-specs.md`
Referenced files: 7
fable-review8.4 KB
---
name: fable-review
description: "Perform an independent, evidence-grounded review of git diffs against requested specifications, architectural invariants, and code standards. Use when reviewing pull requests, inspecting code changes before merge, auditing diffs for regressions, or performing pre-commit sanity reviews — even if the user does not explicitly say \"fable-review\" (e.g. \"review this diff\", \"check this PR\", \"critique my changes\", \"code review this branch\"). Do NOT use for implementing fixes directly (use fable-tdd/fable-execute) or executing tests (use fable-verify)."
version: 1.3.0
pack: proof
inputs:
- implementation_diff
requires:
- target_scope
produces:
- review_evidence
- review_verdict
gates:
- grounded_diff_read
- actionable_findings
fallback: fable-recover
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-verify
continuations:
- fable-security
- fable-release
lateral_peers:
- fable-security
recovery: fable-recover
---
# Fable Review
Review the change as an independent engineer trying to find plausible defects, not as the implementer explaining why the patch is probably fine.
## Mission
A useful review connects a concrete line/change to a concrete failure mode. It prioritizes correctness, invariants, compatibility, lifecycle behavior, and test adequacy before style preference.
The reviewer should be skeptical without manufacturing noise.
## Activate When
- a diff/PR/implementation is ready for independent inspection;
- verification is green but human/semantic risks remain;
- repository conventions or public contracts may have been violated;
- a release/merge needs grounded review evidence.
## Do Not Activate When
- no diff or concrete change exists;
- the main task is automated execution evidence (`fable-verify`);
- the requested work is threat modeling/security specialization (`fable-security`);
- blocking behavior is already known and needs implementation (`fable-execute`).
## Review Classification
Classify the change because different diffs deserve different review depth.
| Change | Primary review focus |
| --- | --- |
| Bug fix | root cause, regression test, adjacent paths |
| New feature | contract, error states, lifecycle, compatibility |
| Refactor | invariant preservation, accidental behavior delta |
| Concurrency | ordering, shared state, cleanup, race/deadlock |
| Persistence/migration | partial failure, transactions, compatibility, rollback |
| Public API/CLI | callers, defaults, error/exit behavior, versioning |
| Dependency upgrade | changed semantics, transitive behavior, config |
| Packaging/build | exports, artifact contents, generated files, runtime entrypoints |
## Review Protocol
### Stage 1 — Reconstruct intent independently
Read:
- user/issue/card acceptance;
- diff against the correct base;
- relevant existing contracts/tests/instructions.
State the intended behavior in your own words before judging the implementation.
### Stage 2 — Read the whole diff, then trace risky changes
Do not review isolated snippets only. Identify:
- public/observable behavior delta;
- state/data-flow delta;
- control-flow/error delta;
- lifecycle/resource delta;
- concurrency delta;
- config/generated/package delta.
Trace important changes into callers/callees where a local diff cannot establish correctness.
### Stage 3 — Check invariants and failure paths
For each material change ask:
- What must remain true before/after?
- What happens on invalid input?
- What happens when dependency call fails/partially succeeds?
- Are cleanup/rollback paths complete?
- Can retries duplicate side effects?
- Can async work outlive ownership/lifecycle?
- Can old/new formats/callers coexist?
### Stage 4 — Check tests as evidence, not decoration
Ask:
- Does a test fail on the pre-fix bug/old behavior where appropriate?
- Does it exercise the real changed boundary?
- Are important negative/error/concurrency/compatibility paths missing?
- Were tests weakened/snapshots blindly updated?
- Could implementation be wrong while tests still pass?
Do not demand tests for trivial static changes when no meaningful behavior is testable.
### Stage 5 — Check scope and maintainability only after correctness
Look for:
- hidden unrelated refactors;
- duplicated logic that creates inconsistent behavior;
- new abstractions whose complexity exceeds need;
- API/config names that misrepresent semantics;
- comments/docs inconsistent with new behavior.
Avoid style-only comments unless repository rules make them blocking or they materially reduce readability/correctness.
### Stage 6 — Calibrate findings
Each finding must include:
- severity: blocking / important / suggestion;
- exact file/line or changed symbol;
- concrete failure scenario;
- why existing evidence does not rule it out;
- minimal repair direction where useful.
If you cannot describe a plausible failure mode, it is probably not a defect finding.
### Stage 7 — Produce verdict
- **APPROVE**: no blocking/important correctness issues found; remaining suggestions are optional.
- **CHANGES_REQUIRED**: at least one grounded issue can cause incorrect behavior, contract violation, or unacceptable risk.
- **INCOMPLETE**: review cannot establish correctness because required context/diff/evidence is missing.
## Decision Rules
- Never approve without reading the actual diff against a known base.
- A passing test suite lowers some risk but does not cancel a code-level defect visible in the diff.
- A suspicious pattern is not a finding until tied to a realistic failure mode.
- Missing test is blocking only when the untested behavior is material and existing evidence cannot cover it.
- For concurrency, reason about interleavings/ownership, not just whether promises are awaited.
- For error handling, trace where the error goes and what state may already have changed.
- For migrations, consider partial execution and mixed-version operation.
- For package/config changes, review the user-installed/runtime artifact path, not just source shape.
- If a blocking issue is narrow and understood, produce one repair card to `fable-execute`; if root cause is uncertain/repeated, route to `fable-recover`.
## Invariants
- Review is independent and read-only.
- Findings are grounded in changed code or directly affected contracts.
- No severity inflation to fill a quota.
- No style preference masquerades as correctness.
- Approval does not claim security proof unless security review actually ran.
- Review verdict covers the actual diff/base inspected.
## Failure Taxonomy
### Rubber stamp
Reviewer relies on tests/author summary and barely reads diff. Re-run full diff review.
### Pattern matching without failure model
Reviewer flags a pattern because it "looks bad" but cannot show impact. Investigate or drop it.
### Local-only review
Change is correct locally but breaks caller/contract/lifecycle. Trace affected boundary.
### Test deference
Reviewer assumes green tests prove all semantics. Inspect test adequacy and changed risk.
### Scope blindness
Unrelated mutation or accidental behavior change hides in a large diff. Compare against card/non-goals.
### Noise overload
Many cosmetic suggestions obscure a real defect. Prioritize by impact and remove quota-driven comments.
## Anti-Patterns
- approving based on PR description alone;
- line-by-line style commentary before understanding behavior;
- "add error handling" without naming a failing error path;
- "add tests" without naming the missing risk;
- flagging every `any`, TODO, or long function independent of change impact;
- assuming an awaited promise means concurrency is safe;
- reviewing only files changed without following a public contract to callers;
- treating absence of findings as evidence the review was deep.
## Finding Template
```text
Severity:
Location:
Changed behavior/invariant:
Failure scenario:
Why current evidence does not cover it:
Suggested bounded repair:
```
## Completion Criteria
Review completes when:
- complete relevant diff/base was read;
- intended behavior and changed risks were reconstructed;
- important invariants/error/concurrency/compatibility/test surfaces were checked as applicable;
- every reported issue has a concrete failure mode and location;
- low-value noise is removed;
- verdict is APPROVE, CHANGES_REQUIRED, or INCOMPLETE with evidence.
## Progressive Resources
- Deep guide: `references/behavioral-diff-review-playbook.md`
- Existing checklist: `references/diff-review-checklist.md`
- Example: `examples/code-review-finding.md`
Referenced files: 7
fable-run8.05 KB
---
name: fable-run
description: "Launch, manage, and verify live applications across CLI binaries, web servers, TUIs, Electron apps, and background daemons with readiness probes and clean teardown. Use when starting development servers, executing live smoke tests, testing interactive CLI binaries, or driving runtime smoke verification — even if the user does not explicitly say \"fable-run\" (e.g. \"start the dev server\", \"launch the app\", \"test the running CLI\", \"smoke test the web app\"). Do NOT use for static analysis or unit testing without a running process (use fable-verify)."
version: 1.3.0
pack: system
inputs:
- app_target
requires:
- built_artifact
produces:
- runtime_evidence
- smoke_proof
gates:
- process_clean_exit
- status_200
fallback: fable-recover
mutatesWorkspace: false
parallelSafe: false
neural_links:
precursors:
- fable-dataviz
- fable-execute
continuations:
- fable-verify
- fable-release
lateral_peers:
- fable-verify
recovery: fable-recover
---
# Fable Run
Execute the real artifact in a bounded environment and collect runtime evidence without confusing "process started" with "application works."
## Mission
Runtime work needs lifecycle ownership: exact command/artifact, working directory, environment, ports, readiness criteria, output bounds, timeout, and cleanup.
A server PID, HTTP 200 from the wrong process, or CLI exit 0 that never exercised the feature is not sufficient proof.
## Activate When
- a server/service must run to verify behavior;
- a built/published CLI or executable needs a smoke test;
- browser/runtime integration requires a live process;
- a runtime path differs materially from unit-test/source execution;
- verification needs empirical entrypoint/lifecycle evidence.
## Do Not Activate When
- unit/integration tests can prove behavior without a persistent live process (`fable-verify`);
- code still needs implementation (`fable-execute`);
- repeated runtime failure/staleness needs diagnosis (`fable-recover`);
- launching the process would cause an unauthorized/destructive external side effect.
## Runtime Classification
| Target | Important controls |
| --- | --- |
| one-shot CLI | cwd/env/args, exit, stdout/stderr, side effects |
| HTTP server | port ownership, readiness, health vs feature probe, cleanup |
| worker/daemon | startup readiness, queue/input fixture, termination |
| browser app | server URL, browser readiness, console/network errors |
| packaged binary | exact artifact/version/path, clean environment |
| multi-service | dependency startup order, ports, teardown, correlation |
## Protocol
### Stage 1 — Identify the exact artifact
Record:
- command/binary path;
- version/hash/build if relevant;
- cwd;
- required environment/config;
- expected process type;
- safe input/probe;
- destructive/external side effects to avoid.
Do not assume `foo` on PATH is the artifact just built—resolve it when identity matters.
### Stage 2 — Establish resource ownership
Before launch determine:
- requested/available port;
- whether an existing process owns it;
- temp directories/files;
- child-process behavior;
- timeout/budget;
- cleanup method.
Never kill an unrelated process merely to acquire a preferred port.
### Stage 3 — Launch with bounded observation
Capture PID/process handle, stdout/stderr, startup errors, and timestamps. Use output/timeout limits so a noisy/hung service cannot consume unbounded resources.
### Stage 4 — Distinguish startup from readiness
A successful spawn only means the OS accepted the process.
Readiness may require:
- port accepting connections;
- health endpoint with expected payload;
- dependency initialization complete;
- CLI command reaches expected branch;
- browser page loads without fatal console/runtime errors.
Use a bounded readiness loop with backoff rather than a fixed arbitrary sleep when possible.
### Stage 5 — Probe the behavior that matters
Health 200 proves health only. If the changed feature is `/checkout`, probe the smallest safe checkout behavior/contract rather than concluding from `/healthz` alone.
Record request/input and relevant response/side effect.
### Stage 6 — Detect wrong/stale process
If output contradicts source/build expectations, check:
- resolved executable/import path;
- process start time;
- port owner PID;
- build/artifact hash;
- cwd/env;
- old server still running.
Route repeated ambiguity to recovery before changing product logic.
### Stage 7 — Terminate cleanly
Stop only processes/resources owned by this run. Wait for clean exit where feasible, escalate termination only within the owned process tree, and remove temporary resources.
### Stage 8 — Record runtime evidence narrowly
State what the run proves: entrypoint launches, endpoint behavior observed, shutdown clean, etc. Hand broader correctness to `fable-verify`.
## Decision Rules
- Port conflict with unknown process → choose another safe port or inspect owner; do not indiscriminately kill it.
- Spawn success without readiness → keep waiting/probing within budget, not PASS.
- Health success without changed-feature probe → only health claim passes.
- Fixed sleep for readiness is weaker than bounded condition polling; prefer observable readiness.
- Runtime output unchanged after source change → prove artifact/process identity before another code mutation.
- One-shot CLI that intentionally exits nonzero can still be a valid tested error path; compare exit/output to expected contract rather than requiring 0 universally.
- External production-like side effect requires explicit safe scope/authorization; prefer local fixture/sandbox when available.
- Background processes must have ownership and cleanup even if the test itself fails.
## Invariants
- Exact runtime artifact/process identity is knowable when used as evidence.
- No unrelated process is killed.
- Spawn, readiness, behavior, and shutdown are separate claims.
- Runtime probes are bounded by time/output/resources.
- Owned background resources are cleaned on success and failure.
- Evidence does not claim more than the actual probe exercised.
## Failure Taxonomy
### Spawn failure
Executable missing, permission/config/startup error. Inspect command/artifact/env before retry.
### Readiness failure
Process alive but service never becomes usable. Capture startup logs/dependencies and diagnose.
### Wrong-process evidence
Port/path points to an older/unrelated process. Resolve identity, do not mutate product.
### Feature probe failure
Runtime is ready but target behavior is wrong. Return to verify/execute/recover based on clarity.
### Hang/leak
Process or child does not terminate. Inspect lifecycle/cleanup; force only owned tree within bounds.
### Environmental mismatch
Behavior depends on cwd/env/OS/runtime version. Record and reconcile rather than treating as product fact universally.
## Anti-Patterns
- `sleep 5 && curl /healthz` as universal runtime proof;
- assuming PID created means ready;
- using 200 from health endpoint to prove unrelated feature;
- killing whatever owns port 3000;
- leaving server running after failure;
- testing a PATH-installed old binary instead of candidate artifact;
- ignoring stderr because process stayed alive;
- unbounded log capture or polling;
- performing real destructive transactions for a smoke test without need.
## Runtime Packet
```text
Artifact/command/version:
CWD/env/config:
Owned resources/PID/port:
Readiness condition/result:
Behavior probe/result:
Stdout/stderr highlights:
Artifact/process identity checks:
Termination/cleanup:
What this proves:
What remains for verify:
```
## Completion Criteria
Runtime execution completes when:
- exact candidate process/artifact was launched under known context;
- readiness was observed, not assumed;
- relevant behavior was probed or evidence scope is explicitly limited;
- stale/wrong process ambiguity is resolved;
- owned processes/resources are cleaned;
- evidence is handed to verification without overclaiming.
## Progressive Resources
- Deep guide: `references/process-lifecycle-and-readiness.md`
- Existing protocol: `references/runtime-process-management.md`
- Example: `examples/live-server-smoke-check.md`
Referenced files: 7
fable-security9.65 KB
---
name: fable-security
description: "Conduct threat modeling, vulnerability assessments, secret sanitization, and security reviews across trust boundaries, auth flows, and untrusted inputs. Use when auditing authentication/authorization logic, inspecting APIs for injection/CORS/CSRF risks, checking for hardcoded credentials, or reviewing security-sensitive diffs — even if the user does not explicitly say \"fable-security\" (e.g. \"security audit this code\", \"check for vulnerabilities\", \"verify auth logic\", \"scan for leaked secrets\"). Do NOT use for general style reviews (use fable-review) or non-security bug fixes (use fable-tdd)."
version: 1.3.0
pack: proof
inputs:
- security_scope
requires:
- threat_boundary
produces:
- security_evidence
- threat_boundary_verdict
gates:
- threat_surface_checked
- no_exposed_secrets
fallback: fable-plan
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-verify
- fable-review
continuations:
- fable-release
lateral_peers:
- fable-review
recovery: fable-recover
---
# Fable Security
Reason about attacker-controlled paths and trust boundaries until every security finding has a concrete source, sink, prerequisite, and impact.
## Mission
Security review is not a secret scan plus a generic checklist. It asks what an attacker can control, which privilege/data boundary that input can cross, what enforcement must hold, and whether the changed code preserves that property under failure and concurrency.
A clean scanner is evidence about scanner coverage, not a proof that the design is secure.
## Activate When
- authentication, authorization, sessions, tokens, permissions, tenancy, or privileged actions change;
- untrusted input reaches parsers, queries, templates, files, URLs, shells, webhooks, or deserializers;
- secrets/credentials or cryptographic material are handled;
- a diff changes trust boundaries, network exposure, storage access, package/install logic, or sandboxing;
- a reported vulnerability/finding needs validation or severity calibration;
- repository/package supply-chain behavior needs security review.
## Do Not Activate When
- the task is only functional verification with no security question (`fable-verify`);
- a generic maintainability review has no trust-boundary impact (`fable-review`);
- a vulnerability is already validated and the user wants only a bounded fix (`fable-execute`, while preserving security acceptance).
## Security Work Classification
| Mode | Primary question |
| --- | --- |
| Threat model | what assets/actors/boundaries/abuse paths exist? |
| Security diff review | what security property changed in this diff? |
| Finding validation | can attacker-controlled data actually reach a sensitive sink? |
| Repository audit | which exposed surfaces deserve deeper inspection? |
| Secret hygiene | can sensitive values enter source/logs/artifacts? |
| Supply-chain/package | can dependency/install/build boundaries be abused? |
## Protocol
### Stage 1 — Define assets, actors, and trust boundaries
Identify:
- protected assets/data/operations;
- authenticated/unauthenticated/privileged actors;
- tenant/user ownership boundaries;
- external systems/webhooks/plugins;
- process/filesystem/network privilege transitions.
Security conclusions without a named boundary are usually too vague.
### Stage 2 — Trace attacker-controlled input to effect
For each relevant entrypoint follow:
`source → parsing/normalization → validation → authorization → transformation → sensitive sink/side effect`
Record where each security property is enforced and whether later transformations can invalidate earlier validation.
### Stage 3 — Check the property appropriate to the boundary
Consider as relevant:
- authentication vs authorization separation;
- object/tenant ownership (IDOR/BOLA);
- CSRF/state binding and redirect validation;
- injection/query/template/shell boundaries;
- path traversal/symlink/archive extraction;
- SSRF and URL/DNS/redirect handling;
- file upload content/size/storage/execution boundaries;
- deserialization/parser resource abuse;
- session/token expiry, replay, rotation, audience/issuer;
- secrets in logs/errors/artifacts;
- fail-open error handling;
- race/TOCTOU and check-then-act authorization;
- privilege escalation via configuration/plugins/hooks;
- dependency/install script/supply-chain assumptions.
### Stage 4 — Model failure and alternate paths
Ask:
- What happens when validation service fails?
- Is denial the default or does code continue?
- Can retries duplicate a privileged action?
- Can an attacker alter state between check and use?
- Do background jobs re-check authorization or trust stale caller claims?
- Does a redirect/proxy/parser transform the value after validation?
### Stage 5 — Validate findings skeptically
For each candidate finding establish:
- attacker prerequisites/control;
- reachable source;
- exact sink/privileged effect;
- missing/bypassed control;
- realistic exploitation path;
- impact/scope;
- existing mitigations that may invalidate or reduce severity.
Do not report a vulnerability merely because a dangerous API exists.
### Stage 6 — Calibrate severity
Severity depends on exploitability + privilege gained + data/tenant scope + required conditions, not scanner labels alone.
Use `blocking` for plausible exploitable issues that violate a required security property. Mark uncertain candidates as needing validation instead of inflating them.
### Stage 7 — Produce bounded remediation evidence
A repair recommendation should name the property to restore and how to prove it, e.g. tenant ownership enforced atomically before mutation, canonical path contained after symlink resolution, redirect allowlist applied to normalized destination.
### Stage 8 — Keep security and functionality separate
Security evidence can block release, but does not substitute for functional test/build/runtime evidence. After repair, both security-specific and functional verification may be required.
## Decision Rules
- Authentication proves identity; it does not prove authorization for a resource/action.
- Validate/authorize as close as practical to the sensitive side effect, especially across async/background boundaries.
- Normalize/canonicalize before containment/allowlist checks when transformations can change meaning.
- Prefer allowlists/capability checks over blocklists for constrained destinations/actions.
- A secret removed from current source may still exist in Git history/logs/artifacts; rotate exposed credentials when exposure is credible.
- Scanner finding with no reachable attacker-controlled path is not automatically exploitable; validate source-to-sink.
- Scanner clean result does not close design-level authz/logic threats.
- Do not execute destructive exploit payloads against real systems; use safe local/fixture proof or code-path reasoning.
- If remediation changes product behavior, hand to TDD/execute and require functional verification afterward.
## Invariants
- Raw secrets are never reproduced in evidence or logs.
- Every reported vulnerability has a concrete failure property and reachable path or is clearly labeled unvalidated.
- Tenant/resource authorization is evaluated independently from login status.
- Security review remains read-only unless explicitly shifted to remediation.
- Functional correctness and security correctness remain separate evidence classes.
## Failure Taxonomy
### False positive
Dangerous primitive exists but attacker cannot control source/reach sink or mitigation blocks it. Downgrade/drop with evidence.
### Hidden trust transition
Code crosses queue/plugin/proxy/background/storage boundary and assumes upstream validation. Trace/revalidate required property.
### Fail-open path
Control error/timeout falls through to privileged behavior. Treat as high-priority boundary failure.
### TOCTOU/race
Authorization/containment check is separated from side effect and mutable state can change. Seek atomic primitive/revalidation.
### Sanitization mismatch
Validation occurs before decode/normalization/redirect/symlink resolution. Check canonical value at sink boundary.
### Secret exposure
Credential entered source/log/artifact/history. Remove exposure path and require rotation/containment as applicable.
## Anti-Patterns
- "no high CVEs, therefore secure";
- generic OWASP checklist without tracing changed paths;
- equating authenticated user with authorized user;
- regex path checks before canonical resolution;
- trusting webhook fields solely because request reached a webhook endpoint;
- reporting theoretical API danger with no attacker path;
- logging raw tokens to debug auth;
- testing exploitability against production without authorization;
- fixing security by disabling validation or broadening permissions to make tests pass.
## Security Finding Packet
```text
Security property:
Asset/boundary:
Attacker prerequisites:
Source → transformations → control → sink:
Concrete failure scenario:
Existing mitigations:
Severity + rationale:
Evidence/location:
Bounded remediation property:
Security re-test:
Functional verification still required:
```
## Completion Criteria
Security work completes when:
- relevant trust boundaries and attacker-controlled paths were traced;
- candidate findings were validated/calibrated rather than pattern-matched;
- secrets were not exposed during analysis;
- blocking findings have bounded remediation/verification criteria;
- verdict states exactly what security property was checked and what remains not checked.
## Progressive Resources
- Deep guide: `references/trust-boundary-and-finding-validation.md`
- Existing threat matrix: `references/threat-modeling-matrix.md`
- Secret hygiene: `references/secret-sanitization.md`
- Example: `examples/security-audit-walkthrough.md`
Referenced files: 8
fable-simplify7.61 KB
---
name: fable-simplify
description: "Refactor and simplify settled, recently modified code to improve readability, remove dead branches, flatten deeply nested logic, and reduce duplication while preserving behavior. Use when cleaning up complex functions, eliminating boilerplate, deduplicating logic, or improving code altitude after tests pass — even if the user does not explicitly say \"fable-simplify\" (e.g. \"simplify this code\", \"clean up this function\", \"refactor this logic\", \"make this cleaner\"). Do NOT use when tests are failing (use fable-recover) or for speculative architectural rewrites (use fable-plan)."
version: 1.3.0
pack: system
inputs:
- target_module
requires:
- passing_tests
produces:
- simplified_diff
gates:
- behavior_preserved
- tests_pass
fallback: fable-verify
mutatesWorkspace: true
parallelSafe: false
neural_links:
precursors:
- fable-execute
continuations:
- fable-verify
- fable-review
lateral_peers:
- fable-execute
recovery: fable-recover
---
# Fable Simplify
Reduce accidental complexity while proving the externally relevant behavior and contracts did not change.
## Mission
Simplification is not permission for a broad rewrite. It should make the code easier to reason about through small, reviewable transformations backed by characterization/invariant evidence.
"Cleaner" is subjective. Behavior preservation, reduced duplication/branching/indirection, and clearer ownership can be demonstrated.
## Activate When
- duplicated logic creates divergence risk;
- nested/control-flow complexity obscures invariants;
- dead code/obsolete indirection is proven unused;
- a refactor is explicitly requested with no intended behavior change;
- a completed implementation needs bounded cleanup before review.
## Do Not Activate When
- user-visible behavior/API/schema is intended to change (`fable-tdd`/`fable-execute`);
- behavior is not characterized well enough to preserve;
- architecture itself is being redesigned (`fable-plan`);
- a cleanup opportunity is unrelated to the active card.
## Simplification Classification
| Smell | Safe first move |
| --- | --- |
| duplicated behavior | prove equivalence, extract shared rule |
| nested branching | identify decision table/invariants, flatten incrementally |
| dead code | prove no runtime/registration/reflection reachability |
| wrapper/indirection | prove callers/semantics before collapse |
| data transformation chain | characterize representative + boundary inputs |
| public/internal API clutter | preserve public contract; simplify behind boundary |
| generated code noise | change generator/source, not output manually |
## Protocol
### Stage 1 — Establish preservation contract
State what must not change:
- public inputs/outputs/errors;
- side effects/state transitions;
- serialization/order/timing where contractual;
- package/API/CLI compatibility;
- performance requirements if material.
Capture baseline tests/fixtures/runtime evidence. If coverage is weak, add characterization before structural mutation.
### Stage 2 — Identify the complexity to remove
Name it concretely: duplicated condition, unnecessary branch, dead adapter, repeated parsing, ownership split, obsolete compatibility path.
Avoid "modernize this module" as an unbounded objective.
### Stage 3 — Prove reachability/deadness before deletion
Search callers plus dynamic registration/plugin/reflection/generated paths. "grep found no import" is not enough for code that may load dynamically.
### Stage 4 — Refactor one semantic step at a time
Examples:
- extract a shared pure rule;
- replace nested branch with explicit guard/decision table;
- collapse wrapper whose contract is identical;
- remove dead branch after reachability proof.
Run the narrow preservation checks after each meaningful step.
### Stage 5 — Watch for accidental contract changes
Review diff for:
- changed exception/error type/message relied on by callers;
- iteration/order changes;
- truthiness/null semantics;
- eager vs lazy evaluation;
- async sequencing;
- transaction/cleanup movement;
- public export/signature/default changes;
- performance/resource behavior if required.
### Stage 6 — Measure simplification honestly
Useful evidence can include:
- fewer duplicated rules/branches;
- smaller public surface;
- fewer states/indirection layers;
- improved named invariant ownership.
Line count alone is not a quality metric.
### Stage 7 — Fresh verify and review
After the last refactor mutation, run fresh affected verification and independent diff review. A pre-refactor green baseline is stale for the simplified code.
## Decision Rules
- No characterization/proof for behavior-rich legacy code → add it before refactor.
- Removing dead code requires runtime reachability confidence, including dynamic registries/plugins.
- If simplification reveals a desired behavior change, split it into a separate TDD/execution card.
- Prefer a few reversible transformations over a from-scratch rewrite.
- Do not collapse abstraction when it represents a real domain/ownership boundary even if it adds lines.
- Avoid DRY when two similar blocks encode different future invariants; shared syntax is not always shared concept.
- If tests break, diagnose whether refactor changed behavior or test was coupled to internals; do not blindly revert/update assertions.
## Invariants
- Accepted external behavior remains unchanged.
- Public contracts do not drift accidentally.
- Each transformation is reviewable and evidence-backed.
- Unrelated cleanup stays out of scope.
- Dynamic/generated reachability is considered before deletion.
- Verification is fresh after final mutation.
## Failure Taxonomy
### Behavior drift
Refactor changes output/error/state/order. Restore invariant or separate intended behavior change.
### Test coupling
Tests fail because they assert internal structure rather than contract. Determine whether test or code should change from accepted behavior—not preference.
### False dead code
Dynamic registration/reflection/plugin path still reaches code. Restore and improve discovery evidence.
### Abstraction collapse
Removed layer encoded a real boundary/invariant. Reintroduce clearer boundary rather than optimizing line count.
### Scope creep
Refactor expands into architecture/feature work. Stop and split/replan.
### Complexity displacement
Code gets shorter locally but moves complexity into callers/config/generic abstraction. Evaluate whole affected decision surface.
## Anti-Patterns
- rewriting a module from scratch to "simplify" it;
- using line count as primary success metric;
- deleting code because no static import exists;
- DRY extraction across semantically different rules;
- changing API signatures during cleanup;
- updating tests to match refactor without checking contract;
- bundling unrelated cleanup after feature work;
- replacing explicit domain logic with clever generic abstraction.
## Simplification Packet
```text
Preservation contract:
Baseline evidence:
Complexity targeted:
Reachability/deadness evidence:
Transformations:
Risky semantic deltas checked:
Fresh verification:
Review result:
Measured simplification:
Deferred behavior changes:
```
## Completion Criteria
Simplification completes when:
- complexity reduction is concrete and scoped;
- behavior/public contracts are characterized and preserved;
- deletion/refactor reachability assumptions are grounded;
- final evidence is fresh;
- independent review finds no hidden behavior drift;
- no unrelated behavior change is smuggled into cleanup.
## Progressive Resources
- Deep guide: `references/behavior-preserving-refactor-strategy.md`
- Existing patterns: `references/refactoring-patterns.md`
- Example: `examples/flatten-nested-logic.md`
Referenced files: 7
fable-simulator7.43 KB
---
name: fable-simulator
description: "Verify complex code changes against independent mathematical oracles, derived specifications, headless browser environments, and isolated sandbox states. Use when deriving independent verification oracles, testing algorithms against reference models, verifying UI with headless browsers, or testing state transitions in sandboxes — even if the user does not explicitly say \"fable-simulator\" (e.g. \"simulate this workflow\", \"verify with an independent oracle\", \"headless browser check\", \"derive a test oracle\"). Do NOT use for simple unit test runs (use fable-verify)."
version: 1.3.0
pack: system
inputs:
- verification_target
requires:
- independent_oracle
produces:
- oracle_evidence
- causal_verification_matrix
gates:
- oracle_independent
- untracked_files_preserved
fallback: fable-verify
mutatesWorkspace: false
parallelSafe: false
neural_links:
precursors:
- fable-verify
continuations:
- fable-verify
- fable-review
lateral_peers:
- fable-verify
recovery: fable-recover
---
# Fable Simulator
Create an independent model of expected behavior and compare it with the candidate without pretending the simulation itself is production evidence.
## Mission
Simulation is useful when ordinary tests risk sharing the same assumptions as the implementation, when a UI/agent flow needs controlled playthrough, or when failure injection can expose paths that are hard to reproduce safely.
The Skill must preserve a hard boundary between **simulated/oracle evidence** and **real runtime verification**.
## Activate When
- an algorithm/rewrite needs an independently derived oracle;
- a contract is implicit across callers and needs reconstruction;
- UI/agent state transitions benefit from scripted playthrough/failure injection;
- testing destructive/rare conditions safely requires a model/sandbox;
- differential testing between two independent implementations is valuable.
## Do Not Activate When
- normal tests/runtime evidence already proves the claim cheaply;
- the proposed oracle is derived from the same code/assumptions under test;
- simulation would be reported as proof that an external production system actually behaved that way.
## Simulation Classification
| Mode | Independence requirement |
| --- | --- |
| Golden fixtures | fixture expected outputs come from trusted contract/observations |
| Reference implementation | independently written/maintained logic |
| Differential provider/model | separate implementation/model with blinded oracle |
| State-machine simulation | transitions/invariants defined from product contract |
| Failure injection | controlled faults with explicit modeled assumptions |
| UI playthrough | browser actions + observable DOM/network/runtime evidence |
## Protocol
### Stage 1 — Define the claim and oracle independence
State what candidate behavior is being checked and why the oracle does not simply repeat candidate logic.
List shared assumptions. If a shared assumption could cause both candidate and oracle to be wrong identically, record that coverage gap.
### Stage 2 — Derive contract from independent sources
Use public interfaces, callers, specs, historical golden outputs, primary docs, or separately maintained reference behavior. Avoid reading implementation details solely to recreate the same algorithm.
### Stage 3 — Build representative and edge corpus
Include normal, boundary, invalid, stateful, adversarial, and failure cases relevant to the claim. Preserve user/untracked workspace files and run in isolated temp/sandbox locations where possible.
### Stage 4 — Execute candidate and oracle separately
Capture inputs, outputs, errors, side effects/state transitions, timing only when required, and tool/environment identity.
### Stage 5 — Compare semantically
Normalize only differences that the contract declares irrelevant. Do not normalize away a real behavioral difference just to reach zero diff.
For UI, compare causal rows:
`action → expected observable → actual observable → evidence`
not pixels alone unless pixel fidelity is itself the contract.
### Stage 6 — Triage divergence
A mismatch means one of:
- candidate wrong;
- oracle wrong/stale;
- contract ambiguous;
- normalization invalid;
- environment differs.
Do not automatically "fix candidate to oracle" until the source of truth is established.
### Stage 7 — Mark evidence scope
Report simulation/oracle evidence distinctly. Hand off to `fable-verify` for real runtime/package/environment proof where the claim requires it.
## Decision Rules
- An oracle copied from candidate implementation is self-confirming and invalid.
- A second LLM is not automatically an independent oracle if it receives the candidate answer or hidden expected output.
- Golden fixtures need provenance; unexplained fixtures can encode old bugs.
- UI screenshots without interaction/state evidence are weak for functional causality.
- Failure injection proves behavior under the modeled fault, not that real infrastructure fails exactly that way.
- Zero diff across a narrow corpus is not universal correctness; state coverage explicitly.
- Never modify/delete untracked user files to make simulation deterministic.
## Invariants
- Oracle independence and shared assumptions are explicit.
- Simulation evidence is labeled as simulation.
- Input corpus and normalization are reproducible.
- User/untracked workspace state is preserved.
- Divergence is diagnosed before deciding which side is wrong.
- Real-world claims receive real verification when required.
## Failure Taxonomy
### Tautological oracle
Oracle reuses candidate implementation/answer. Redesign independently.
### Stale oracle
Reference behavior no longer matches accepted contract. Re-ground oracle before judging candidate.
### Ambiguous contract
Candidate and oracle differ but both are plausible. Route to plan/research rather than choose arbitrarily.
### Over-normalization
Comparison strips a meaningful difference. Narrow normalization to contract-declared irrelevant fields.
### Simulation/reality confusion
Simulated pass is used as production/runtime proof. Downgrade claim and hand to verify/run.
### Workspace contamination
Simulation writes into real user state. Move to isolated sandbox and restore owned mutations only.
## Anti-Patterns
- implementing the oracle by copying candidate code;
- giving an LLM oracle the expected answer;
- demanding 100% zero diff without representative corpus reasoning;
- treating screenshots as complete UI verification;
- normalizing every mismatch until green;
- claiming external service behavior from a fake;
- deleting untracked files to reset simulation;
- changing candidate immediately on any oracle mismatch.
## Simulation Packet
```text
Claim:
Oracle type/source:
Independence + shared assumptions:
Corpus/failure injections:
Candidate environment:
Comparison/normalization:
Divergences:
Oracle validity decision:
Simulation verdict:
Real evidence still required:
```
## Completion Criteria
Simulation completes when:
- oracle independence is credible and documented;
- corpus covers the important contract dimensions;
- candidate/oracle comparisons are reproducible;
- divergence is classified instead of blindly resolved;
- workspace remains safe;
- evidence is labeled narrowly and handed to real verification where needed.
## Progressive Resources
- Deep guide: `references/oracle-independence-and-simulation-boundaries.md`
- Existing guide: `references/oracle-derivation-guide.md`
- Example: `examples/independent-oracle-verification.md`
Referenced files: 7
fable-skill-creator11.2 KB
---
name: fable-skill-creator
description: "Author, evaluate, refine, optimize, and package autonomous AI agent skills across multi-agent ecosystems with BinEval scoring and description tuning. Use when creating a new skill from scratch, editing an existing skill that misfires or undertriggers, authoring test suites for skills, optimizing skill descriptions, or packaging skills for distribution — even if the user does not explicitly say \"fable-skill-creator\" (e.g. \"create a skill\", \"build a new skill\", \"optimize skill description\", \"package this skill\", \"teach the agent to do X\"). Do NOT use for general application code changes or non-skill tasks."
version: 1.3.0
pack: creator
inputs:
- user_intent
- workflow_trace
- reference_sources
requires:
- clear_capability_scope
produces:
- structured_skill_package
- eval_benchmark_suite
gates:
- lack_of_surprise
- objective_eval_criteria
fallback: fable-plan
mutatesWorkspace: true
parallelSafe: false
neural_links:
precursors:
- fable-discover
- fable-research
continuations:
- fable-eval
- fable-verify
- fable-review
lateral_peers:
- fable-artifact
- fable-config
recovery: fable-recover
---
# Skill Creator
Create Skills that change agent behavior under pressure, not Skills that merely describe good practice.
## Mission
A production Skill is a compact operating manual for a recurring class of decisions. It should help an agent recognize the situation, choose among plausible actions, collect the right evidence, reject tempting shortcuts, and hand off cleanly when another specialist is better suited.
A schema-valid `SKILL.md` is not enough. A Skill that only says "inspect, plan, implement, verify" is documentation, not a behavioral capability.
## Activation Contract
Use this Skill when the requested work is to:
- create a new reusable Skill;
- strengthen or refactor an existing Skill;
- improve triggering accuracy or reduce overlap with neighboring Skills;
- add references, examples, templates, agents, or evals;
- convert a repeated workflow into a portable agent capability;
- prove that a Skill behaves correctly across more than a happy-path example.
Do not use it for ordinary application changes, bug fixes, or repository maintenance. Route those to the appropriate lifecycle Skill.
## Inputs
- **`user_intent`**: capability, behavior, or workflow to encode.
- **`workflow_trace`**: observed sequence of decisions/tools/failures, when available.
- **`reference_sources`**: primary docs, schemas, domain rules, or proven internal patterns.
## Skill Depth Classification
Before authoring, classify the capability:
| Class | Typical shape | Required depth |
| --- | --- | --- |
| Rule | One stable decision with few branches | concise Skill, strong boundaries |
| Procedure | Multi-step repeatable operation | staged protocol + evidence |
| Diagnostic | Symptoms map to competing causes | failure taxonomy + falsification |
| Orchestrator | Chooses/delegates among specialists | routing table + handoff contracts |
| Domain expert | Requires substantial specialist knowledge | deep references + examples + eval families |
If the capability is Diagnostic, Orchestrator, or Domain expert, a short checklist is presumptively insufficient.
## Skill Contract V2
Every non-trivial Skill must answer these questions explicitly.
### 1. When does it activate?
Define positive triggers, negative triggers, prerequisites, and escalation conditions. Include near-neighbor tasks that should *not* select this Skill.
### 2. What situation is the agent in?
Provide a small classification that changes behavior. Examples: bounded vs architectural, deterministic vs flaky, local vs external, reversible vs destructive, known vs uncertain.
### 3. What decisions must be made?
Use decision tables or branching rules for ambiguous cases. Do not hide judgment behind vague language such as "when appropriate" without defining the evidence that makes it appropriate.
### 4. What is the execution protocol?
For each stage specify:
- objective;
- allowed actions;
- required evidence;
- exit condition;
- stop/escalation condition.
### 5. How can it fail?
Name domain-specific failure classes and observable signals that distinguish them. Do not collapse every failure into "retry" or "use recover".
### 6. What must never become false?
List invariants that stay true throughout the Skill. Examples: no production mutation before valid RED, no release claim without registry confirmation, no review finding without a concrete failure mode.
### 7. Which shortcuts are tempting but invalid?
Document anti-patterns that capable models commonly choose under time pressure or ambiguous prompts.
### 8. What evidence leaves the Skill?
Define a handoff packet for the next specialist: facts learned, commands run, artifacts changed, residual risk, unresolved questions, and proof freshness where relevant.
### 9. What belongs in progressive resources?
`SKILL.md` should hold decision logic. References should add operational depth, not repeat the same checklist in different words.
Useful resources include:
- decision heuristics;
- failure diagnosis guides;
- legacy/edge-case playbooks;
- worked examples with trade-offs;
- tool-independent verification strategies;
- checklists only where the checklist encodes non-obvious domain knowledge.
### 10. How will we know the Skill actually works?
Create multiple semantically distinct eval families. Surface variations of one scenario do not count as breadth.
## Eval Depth Standard
For a non-trivial Skill, target at least **6 semantic scenario families** across its package before calling evaluation coverage mature. A family is a different decision problem, not the same prompt rewritten five ways.
Include a mix of:
- straightforward activation;
- boundary/non-trigger case;
- ambiguous situation requiring classification;
- adversarial pressure to take a shortcut;
- partial or contradictory evidence;
- failure/recovery handoff;
- legacy or constrained environment where the normal happy path is unavailable.
The enterprise behavior harness may still expand each family into known/negative/ambiguous/adversarial/holdout variants. That expansion does not replace semantic family breadth.
## Authoring Procedure
1. **Observe the real job**: identify decisions, artifacts, failure modes, and handoffs from real workflows or primary sources.
2. **Map neighboring Skills**: define what this Skill owns and what it deliberately refuses.
3. **Write eval scenarios first** for the most important failure-prone decisions. If the expected action cannot be stated clearly, the Skill contract is not ready.
4. **Design the decision model**: classification, branches, invariants, stop conditions.
5. **Write `SKILL.md`** in imperative, tool-independent language.
6. **Add progressive resources** for depth that would otherwise make the main Skill noisy.
7. **Add worked examples** including at least one case where the obvious first move is wrong.
8. **Create/refresh agent profiles** so host-facing prompts point to the full Skill contract rather than replacing it with a one-line summary.
9. **Validate packaging and registry metadata**.
10. **Run deterministic and behavioral evals**. Treat changed Skill/eval corpus hashes as requiring fresh proof.
## Invariants
- Never lower evaluation thresholds to preserve a maturity label after a Skill changes.
- Never copy private expected or forbidden oracle fields into provider-visible prompts.
- Canonical Skills must describe capabilities in provider-neutral language.
- Every non-trivial Skill package must include at least one substantial progressive reference (>=1000 bytes).
- Behavioral maturity claims require fresh verified holdout evaluation across semantic scenario families.
## Decision Rules
- Prefer one strong Skill with a clear domain over several overlapping micro-Skills.
- If two Skills can both reasonably trigger from the same ordinary request, sharpen their boundaries before adding more keywords.
- If the body contains only a linear happy-path procedure, add the branches that real work forces an expert to choose between.
- If a reference merely paraphrases the body, replace it with deeper material.
- If an eval can be passed by parroting the Skill description without understanding the situation, add ambiguity or conflicting evidence.
- Never lower eval thresholds to preserve a maturity label after the Skill changes.
- Never copy private expected/forbidden oracle fields into provider-visible prompts.
- Avoid provider-specific tool names in canonical behavior. Describe capabilities, then let host adapters map them.
## Failure Taxonomy
### Triggering failure
Skill activates too often or misses obvious requests. Fix boundaries and description before adding more procedure text.
### Behavioral shallowness
Skill selects the right broad action but cannot discriminate difficult subcases. Add classification, decision branches, and semantic eval families.
### Resource shallowness
References exist but contain only summaries. Replace with operational guides and worked trade-offs.
### Evaluation illusion
Pass rate is high because scenarios are near-duplicates or leak the answer. Increase semantic breadth and oracle isolation.
### Contract drift
`SKILL.md`, package metadata, agent profiles, registry entries, and evals describe different behavior. Reconcile before release.
## Anti-Patterns
Reject these patterns during Skill review:
- "Use best practices" without naming the decision or evidence.
- one generic procedure reused across unrelated Skills;
- references under ~one screen that only restate the main steps;
- one toy example presented as domain coverage;
- an eval suite made from one scenario plus wording variants;
- arbitrary retry counts with no diagnostic reason;
- completion criteria that can be satisfied without fresh evidence;
- `When NOT to Use` sections that omit the closest competing Skill;
- giant monolithic prompts that duplicate every reference and destroy progressive disclosure.
## Evidence Requirements
A strengthened Skill should provide:
- valid package/frontmatter structure;
- explicit activation and refusal boundaries;
- situation classification and decision rules;
- domain-specific failure taxonomy;
- at least one substantial progressive reference;
- worked example(s) covering a non-happy path;
- multiple semantic eval families;
- behavioral evidence re-run when the evaluated corpus changes.
## Handoff Contract
When finishing Skill authoring, report:
- Skill ID and owned capability;
- neighboring Skills and trigger boundaries;
- semantic eval families added;
- progressive resources added;
- known unsupported cases;
- deterministic validation status;
- whether behavioral evidence is fresh, stale, or not yet executed.
## Completion Criteria
Do not call the Skill mature merely because all expected folders exist.
Completion requires:
- package validation passes;
- contract sections above are materially present where relevant;
- resources add depth rather than duplication;
- eval breadth tests the major decisions and shortcuts;
- registry/host metadata remains consistent;
- any prior behavioral proof invalidated by corpus changes is explicitly marked stale until rerun.
## Progressive Resources
- Depth standard: `references/depth-standard-v2.md`
- Existing authoring references and templates remain useful for package mechanics.
Referenced files: 13
fable-spark7.25 KB
---
name: fable-spark
description: "Predict the smallest atomic next engineering action from current workspace state, evidence gates, and mutation freshness with situational silence. Use when evaluating what immediate atomic step to take next, checking situational awareness, or determining if the current state requires action or silence — even if the user does not explicitly say \"fable-spark\" (e.g. \"what is my next move\", \"spark hint\", \"check next action\", \"what should I do now\"). Do NOT use as a substitute for broad architectural planning (use fable-plan)."
version: 1.3.0
pack: system
inputs:
- current_state
requires:
- situational_context
produces:
- spark_suggestion
gates:
- minimal_action_tested
fallback: get-fable
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- get-fable
- fable-cowork
continuations:
- get-fable
lateral_peers:
- get-fable
recovery: fable-recover
---
# Fable Spark
Choose one small next move that improves the state of the current task—or deliberately say nothing.
## Mission
Spark is not a planner, router, motivational assistant, or todo generator. It is a micro-policy for moments when the task is already in motion and the best next action should be small, concrete, and justified by current state.
The quality bar is not "a helpful suggestion." It is **the least action that unlocks the next piece of evidence or safely advances the active Skill**.
## Activate When
- an edit/test/failure/state transition just happened;
- a long session needs the next atomic step rather than another full plan;
- verification freshness/gates determine what should happen next;
- an active specialist has a clear ownership boundary but the immediate move is unclear;
- silence vs intervention itself is a useful decision.
## Do Not Activate When
- no primary Skill has been selected for substantial new work (`get-fable`);
- architecture/decomposition is the unresolved problem (`fable-plan`);
- repeated failure requires a diagnosis (`fable-recover` owns the work; Spark may only point there);
- the user asked for execution, not a suggestion, and the executing Skill already knows its next step.
## Situation Classification
| State signal | Spark posture |
| --- | --- |
| repeated failure | stop mutation; diagnose |
| mutation newer than verification | refresh relevant proof |
| explicit active gate missing | satisfy that gate |
| review finding unresolved | repair/replan finding |
| active card blocked by unknown | discover/research/plan |
| complete/idle with no intent | silence |
| specialist owns next obvious step | usually silence; avoid narration |
| several equally plausible non-atomic moves | defer to orchestrator/plan |
## Next-Move Protocol
### Stage 1 — Read state before intent embellishment
Inspect phase, current Skill/card, mutation/verified generation, failure streak, evidence, unresolved findings, blockers, and latest user intent.
### Stage 2 — Apply safety precedence
Prefer in order when applicable:
1. diagnose repeated/contradictory failure;
2. refresh stale proof after mutation;
3. satisfy explicit blocking gate;
4. resolve a named load-bearing unknown;
5. continue the active specialist's smallest next action;
6. silence.
### Stage 3 — Make the action atomic
A Spark suggestion should usually be one observable verb-object step:
- `run the affected auth tests`;
- `reproduce the race with a barrier`;
- `inspect the registered CLI entrypoint`;
- `review the current diff against the card`;
- `capture the provider response bundle`.
Avoid multi-stage suggestions such as "research, plan, implement, test, and release."
### Stage 4 — Require a reason tied to evidence
The reason should reference state/evidence, e.g. verification stale after mutation, failure streak reached recovery threshold, release gate missing, or active card has unresolved API contract.
### Stage 5 — Check actionability
Before speaking, ask:
- can the action be done now with available context/tools?
- does it preserve current scope?
- is it owned by the active/next Skill?
- will its outcome reduce uncertainty or satisfy a gate?
If not, prefer silence or route back to the orchestrator.
### Stage 6 — Use silence intentionally
Silence is correct when:
- task is idle/complete with no new intent;
- the active agent already has an obvious immediate action;
- only speculative future work can be suggested;
- available context is too ambiguous for a useful atomic step.
## Decision Rules
- Repeated failure outranks "run tests again" if another identical run adds no information.
- Fresh mutation outranks completion/release suggestions until relevant evidence is refreshed.
- Security/research/receipt evidence cannot stand in for required functional verification.
- Do not suggest implementation when a load-bearing fact/contract remains unknown.
- Prefer a targeted affected check over a full suite when the next goal is rapid causal feedback; broader required gates can follow.
- Never invent a new objective because the current task is quiet.
- Confidence is not a license to act outside the current Skill's scope.
- If two next moves depend on an unresolved ordering/architecture choice, route to plan instead of guessing.
## Invariants
- At most one primary Spark suggestion.
- Suggestion is atomic, actionable, and scope-preserving.
- State/evidence—not generic best practice—justifies intervention.
- Silence is allowed and preferred over speculative advice.
- Spark never upgrades stale/incomplete evidence into completion proof.
## Failure Taxonomy
### Scope drift
Suggestion introduces an unrelated improvement. Drop it and return to active card/gate.
### Ceremony
Spark narrates a step the active specialist is already obviously executing. Stay silent.
### Premature action
Suggestion mutates code before discovery/TDD/recovery requirement. Apply precedence.
### Stale-proof blindness
Spark suggests review/release despite mutation newer than verification. Refresh proof first.
### Retry loop
Spark repeats a command after repeated failure without new diagnostic state. Route to recovery.
### Ambiguous macro-action
Suggestion contains several dependent steps. Reduce to first atomic decision or route to plan.
## Anti-Patterns
- turning Spark into a mini project manager;
- always suggesting something;
- generic advice such as "keep testing";
- proposing a full suite when one focused probe is the causal next move;
- suggesting release because tests once passed;
- suggesting implementation from unknown external facts;
- confidence scores unsupported by state;
- outputting three alternatives instead of one next action.
## Spark Packet
```text
State signal:
Blocking gate / uncertainty:
Suggestion: <one action or null>
Why now:
Expected observation:
Owner Skill:
Silent: true|false
```
## Completion Criteria
Spark succeeds when it either:
- names one immediately executable, evidence-grounded action that safely advances the current lifecycle; or
- remains silent because intervention would add noise, scope, or speculation.
## Progressive Resources
- Deep guide: `references/atomic-action-and-silence.md`
- Next-move policy: `references/next-move-policy.md`
- Silence: `references/silence-policy.md`
- Evidence/gates: `references/evidence-and-gates.md`
- Confidence: `references/confidence-policy.md`
- Example: `examples/situational-awareness-walkthrough.md`
Referenced files: 10
fable-tdd9.73 KB
---
name: fable-tdd
description: "Drive testable behavior changes and bug fixes through disciplined red-green-refactor cycles with observable regression tests. Use when implementing new features with unit/integration tests, fixing reproducible bugs, modifying business logic, or writing test-first behavior contracts — even if the user does not explicitly say \"fable-tdd\" (e.g. \"fix this bug test-first\", \"write a test and make it pass\", \"add this feature with tests\", \"TDD this logic\"). Do NOT use for broad exploratory prototyping without clear assertions (use fable-discover or fable-plan) or for post-implementation reviews (use fable-review)."
version: 1.3.0
pack: build
inputs:
- behavior_contract
requires:
- test_harness
produces:
- regression_test
- behavior_change
gates:
- red_observed
- green_observed
fallback: fable-recover
mutatesWorkspace: true
parallelSafe: false
neural_links:
precursors:
- fable-plan
continuations:
- fable-execute
- fable-verify
lateral_peers:
- fable-execute
recovery: fable-recover
---
# Fable TDD
Prove the behavior is missing or broken before changing production code, then make the smallest change that satisfies the right test at the right level.
## Mission
TDD is not "write any failing test first." The red state must demonstrate the intended behavior gap through a trustworthy harness. A syntax error, broken fixture, stale build, or mock-only expectation does not earn the right to change production code.
The Skill optimizes for three things:
- **causal confidence**: the test fails because the behavior is wrong;
- **minimal intervention**: implementation changes only what the behavior requires;
- **durable regression proof**: the test would catch the bug if it returned.
## Activate When
- fixing a reproducible bug or regression;
- adding behavior with a stable enough contract to assert;
- changing validation, state transitions, calculations, API behavior, persistence, or integration semantics;
- refactoring where an invariant needs executable characterization first.
## Do Not Activate When
- the behavior cannot yet be located or reproduced (`fable-discover`);
- the main uncertainty is external API semantics (`fable-research`);
- the change is non-executable docs/metadata with no meaningful behavior test (`fable-execute`);
- the required test harness is itself broken or executing stale artifacts (`fable-recover`).
## Change Classification
Classify the behavior before choosing a test.
| Class | Preferred proof |
| --- | --- |
| Pure/domain logic | focused unit/property test |
| Boundary validation/error mapping | unit or contract test at boundary |
| Cross-module interaction | integration test through the changed contract |
| Database/queue/cache behavior | integration test with realistic boundary where feasible |
| HTTP/CLI/public API | contract/integration test at public entry point |
| UI user flow | component/integration first; E2E for high-value cross-boundary behavior |
| Concurrency/timing | deterministic coordination test, not sleep-and-hope |
| Legacy behavior with poor seams | characterization test at nearest stable boundary |
| Refactor/no intended behavior change | characterization/invariant tests before movement |
Use the lowest test level that proves the real behavior **without mocking away the thing under test**.
## Protocol
### Stage 1 — Write the behavior contract
State:
- initial state/input;
- trigger/action;
- expected observable result;
- relevant side effects;
- error/boundary behavior;
- invariant that must remain true.
For a bug, capture the concrete reproduction separately from the proposed implementation.
### Stage 2 — Validate the harness
Before RED, confirm the chosen test actually executes the relevant path.
Check when applicable:
- source vs built artifact;
- test discovery/config;
- fixture realism;
- feature flags/env;
- mock boundaries;
- asynchronous completion;
- cleanup/isolation;
- whether the assertion observes public behavior rather than an internal call count.
### Stage 3 — Choose the test level
Ask:
1. What production change would make this test fail again?
2. Does the test cross the contract where the bug actually lives?
3. Have I mocked the suspected failure away?
4. Can a narrower test prove the same user-visible behavior more deterministically?
If you cannot answer #1, the test is probably weak.
### Stage 4 — RED
Write the smallest test that captures the behavior contract and run it.
RED is valid only if:
- test executes;
- failure is deterministic enough to reason about;
- failure message/state matches the expected missing/broken behavior;
- it is not failing because of setup, syntax, import, environment, stale artifact, or unrelated test pollution.
If the failure reason is wrong, repair the test/harness and repeat RED. **Do not touch production code yet.**
### Stage 5 — Minimal GREEN
Change the smallest production surface that can satisfy the valid RED.
Do not:
- redesign adjacent APIs;
- add speculative flexibility;
- refactor unrelated code;
- weaken the assertion;
- replace real behavior with a mock to reach green.
Run the focused test until GREEN.
### Stage 6 — Adjacent falsification
Before refactor, probe the closest failure surfaces appropriate to the change:
- boundaries/empty/invalid input;
- error path;
- state transition ordering;
- idempotency/retry;
- concurrency/race;
- compatibility with prior format/API;
- persistence/transaction behavior.
Add tests only when they protect a meaningful contract; do not inflate count mechanically.
### Stage 7 — Refactor under green
Now improve structure if needed. Refactor in small steps and rerun the affected tests after each meaningful mutation.
### Stage 8 — Fresh handoff
Record final mutation and hand off to `fable-verify` with:
- behavior contract;
- exact RED evidence/reason;
- GREEN command/result;
- tests added/changed;
- production surfaces changed;
- adjacent cases probed;
- residual risks not covered by the focused test.
## Decision Rules
- Test passes before production change → it does not prove the requested gap; redesign the test or confirm behavior already exists.
- Test errors before assertion → fix harness/test, not production.
- Bug is nondeterministic → first control or instrument the nondeterminism; repeated random reruns are not a reliable RED.
- Concurrency bug → prefer barriers/latches/fake clocks/deterministic scheduling over arbitrary sleeps.
- Legacy code has no unit seam → test the nearest stable public boundary before introducing a seam; do not perform a broad refactor just to make unit testing aesthetically pure.
- External dependency cannot be exercised locally → use a contract/fake only after establishing what behavior the fake must preserve from primary evidence.
- User asks to "just patch it" → if executable behavior is changing, preserve RED discipline unless the absence of a viable harness is explicit and accepted.
- Test expectation conflicts with current agreed product contract → do not force implementation to an obsolete test; resolve the contract first.
- A test only asserts that a mock was called → add observable outcome/state evidence unless the call itself is the public contract.
## Invariants
- No production behavior mutation before a valid RED for testable changes.
- RED and GREEN refer to the same behavior contract.
- Tests are not weakened to make implementation pass.
- The suspected failure mechanism is not mocked away.
- Final GREEN is fresh after the last relevant mutation.
- Refactoring does not add behavior outside the accepted card.
## Failure Taxonomy
### Wrong RED
Failure comes from syntax/import/setup/fixture/environment rather than target behavior. Repair harness first.
### False GREEN
Test passes but does not cross the real failure boundary, often because a mock or stale artifact bypasses it. Strengthen/reposition the test.
### Flaky RED/GREEN
Outcome changes without relevant code mutation. Identify nondeterminism before treating either state as evidence.
### Untestable legacy surface
No narrow seam exists. Characterize at a stable boundary, then introduce the smallest seam supported by the test.
### Contract ambiguity
Expected behavior itself is disputed/unclear. Return to planning/research rather than encoding a guess as a test.
### Implementation loop
Valid RED exists but two materially similar fixes fail. Stop editing and route to `fable-recover` with the RED, attempts, and observed failure differences.
## Anti-Patterns
- writing implementation, then backfilling a test that immediately passes;
- accepting any failure as RED;
- changing expected output to match current implementation;
- asserting only mock interactions when user-visible state can be asserted;
- using sleeps to "test" a race;
- building an elaborate test abstraction before proving one behavior;
- broad legacy refactor before characterization;
- forcing unit tests where only an integration boundary can prove the behavior;
- skipping fresh GREEN after refactor;
- equating coverage percentage with regression proof.
## Evidence Packet
```text
Behavior contract:
Test level + why:
RED command/result + expected failure reason:
Production mutation:
GREEN command/result:
Adjacent cases probed:
Refactor mutations:
Final fresh GREEN:
Residual risk / verify next:
```
## Completion Criteria
TDD completes when:
- the behavior gap was demonstrated by a valid RED;
- production code changed only after that RED;
- focused GREEN is observed for the same contract;
- meaningful adjacent failure surfaces were considered;
- final tests are fresh after refactor/last mutation;
- `fable-verify` receives enough evidence to independently falsify the result.
## Progressive Resources
- Deep strategy: `references/test-strategy-and-hard-cases.md`
- Existing cycle reference: `references/red-green-refactor.md`
- Example: `examples/failing-test-first.md`
Referenced files: 7
fable-verify8.65 KB
---
name: fable-verify
description: "Falsify software implementations and gather fresh, machine-checked acceptance proof across tests, builds, typechecks, and runtime smoke checks before completion. Use when running test suites, checking type correctness, validating acceptance criteria, or verifying post-mutation code integrity — even if the user does not explicitly say \"fable-verify\" (e.g. \"verify my changes\", \"run all tests and typecheck\", \"check if everything passes\", \"prove this works\"). Do NOT use for planning (use fable-plan) or diff code review (use fable-review)."
version: 1.3.0
pack: core
inputs:
- implementation_diff
requires:
- test_suite
produces:
- verification_evidence
- falsification_verdict
gates:
- fresh_mutation_covered
- machine_checked
fallback: fable-recover
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors:
- fable-execute
- fable-tdd
continuations:
- fable-review
- fable-security
- fable-release
lateral_peers:
- fable-simulator
- fable-run
recovery: fable-recover
---
# Fable Verify
Try to prove the implementation wrong, then report only the claims that survive fresh evidence.
## Mission
Verification is not "run the test suite." It is coverage of changed risk with evidence that is both relevant and fresh.
A green command proves only what it actually exercised. Build success does not prove runtime behavior. Security success does not prove functional correctness. Unit success does not prove a packaging or integration path. Old evidence does not prove a newer mutation.
## Activate When
- implementation or repair is ready for independent proof;
- a completion claim needs fresh evidence;
- a review/release gate requires test/build/runtime evidence;
- stale evidence must be refreshed after mutation;
- the suspected regression surface spans more than the focused TDD test.
## Do Not Activate When
- writing the implementation (`fable-execute`/`fable-tdd`);
- deciding architecture (`fable-plan`);
- reviewing design/maintainability from the diff (`fable-review`);
- diagnosing repeated confusing failures (`fable-recover`).
## Risk Classification
Map each changed surface to the evidence capable of falsifying it.
| Changed surface | Typical evidence |
| --- | --- |
| pure behavior | focused unit/property + affected suite |
| cross-module contract | integration/contract test |
| public CLI/API | invocation/smoke + contract tests |
| build/export/package | build + package/clean-install smoke |
| persistence/migration | integration + migration/compatibility fixtures |
| async/concurrency | deterministic ordering/stress supplement |
| config/feature flag | tests under relevant config branches |
| browser/UI | component/integration/E2E as appropriate |
| security boundary | security-specific checks **plus** functional evidence |
| performance-sensitive path | targeted measurement when requirement exists |
## Verification Protocol
### Stage 1 — Read the diff and execution packet
Do not choose commands from habit alone. Identify:
- behavior changed;
- files/contracts touched;
- tests added/changed;
- generated/package/config surfaces;
- residual risks from implementation;
- current mutation generation.
### Stage 2 — Build a verification matrix
For each material risk, write:
- claim;
- failure mode;
- evidence/command that would catch it;
- whether evidence must be narrow, integration, runtime, build, E2E, security, or package-level.
Remove duplicate checks that prove the same narrow fact; add missing checks for untested surfaces.
### Stage 3 — Run narrow, causal checks first
Start with the focused changed behavior and affected tests. Fast local evidence helps distinguish implementation failure from unrelated suite noise.
### Stage 4 — Expand according to blast radius
Run broader typecheck/build/test/integration/E2E/package gates only where the change can affect them or where repository release policy requires them.
Do not skip required project-wide gates merely because narrow tests pass.
### Stage 5 — Adversarial falsification
Actively probe the most plausible regression surfaces:
- boundary/empty/invalid inputs;
- error propagation;
- retries/idempotency;
- compatibility old/new formats;
- async ordering/concurrency;
- cleanup/resource lifecycle;
- package/export/runtime entry point;
- configuration branch.
Prefer tests/commands that can fail meaningfully over speculative prose.
### Stage 6 — Detect nondeterminism and stale execution
If identical commands alternate outcomes, stop counting passes. Capture the variability and route to recovery.
If results do not reflect known source changes, check source-vs-build, cache, branch, env, process, and artifact path before accepting output.
### Stage 7 — Record typed, fresh evidence
Evidence should include:
- kind (`test`, `build`, `runtime`, `review`, `security` where applicable);
- exact command/probe;
- exit/result and relevant counts;
- mutation generation/artifact SHA when available;
- scope/claim the evidence proves.
### Stage 8 — State the verdict narrowly
Use:
- **PASS**: every required claim has fresh relevant evidence;
- **FAIL**: at least one required claim is falsified;
- **INCOMPLETE**: required evidence cannot be obtained or does not cover the risk.
Never convert INCOMPLETE to PASS because the remaining check is inconvenient.
## Decision Rules
- Evidence generation older than current relevant mutation → stale, rerun.
- Test passes but does not execute changed path → irrelevant evidence, not PASS.
- Security-only evidence for functional bug → require functional proof.
- Build/typecheck-only evidence for runtime behavior → require runtime/test proof.
- Focused test green but integration contract changed → run integration/contract evidence.
- Package/export change → inspect/package/smoke the distributed artifact, not only source tests.
- Flaky alternating outcomes → route to `fable-recover`; do not cherry-pick a green run.
- New verification failure caused by a bounded obvious implementation defect → return one repair card; repeated/ambiguous failure → recover.
- If a command is unavailable in current environment, mark evidence incomplete and name the external gate instead of fabricating a result.
## Invariants
- Verification is read-only with respect to product behavior; any repair starts a new mutation/evidence cycle.
- Every completion claim maps to fresh evidence.
- Evidence kinds are not interchangeable.
- The final verdict covers the current mutation generation/artifact.
- Required failures/warnings are not hidden by truncating output or rerunning until green.
## Failure Taxonomy
### Relevant test failure
Changed behavior is falsified. Return bounded repair or recover depending on clarity/repetition.
### Unrelated suite failure
Prove it is unrelated before excluding it; do not dismiss merely because it predates the change.
### Stale evidence/artifact
Output corresponds to older mutation/build. Refresh build/path/env before evaluating implementation.
### Nondeterministic verification
Same state produces inconsistent result. Diagnose shared state/timing/environment before verdict.
### Coverage gap
Available checks do not exercise material changed risk. Add/locate suitable evidence or mark incomplete.
### Environment limitation
Required external service/browser/platform/release environment unavailable. Report exact missing evidence; do not emulate a pass.
## Anti-Patterns
- `tests passed` as the entire verification report;
- running only tests the implementer just wrote;
- rerunning flaky tests until green;
- trusting old screenshots/logs after code changed;
- treating typecheck/build/security as functional proof;
- ignoring packaging/export paths;
- broad full-suite runs with no affected-risk reasoning;
- claiming pass when an external required gate was never run.
## Verification Packet
```text
Current mutation/artifact:
Changed risks:
Verification matrix:
- claim → evidence → result
Adversarial probes:
Stale/nondeterministic evidence detected:
Required external evidence not run:
Verdict: PASS | FAIL | INCOMPLETE
Repair/recovery recommendation:
```
## Completion Criteria
Verification completes when:
- changed risks have relevant checks;
- required project gates are executed or explicitly marked unavailable;
- evidence is fresh for current mutation/artifact;
- likely edge/failure surfaces were actively probed;
- nondeterminism/stale output is not hidden;
- verdict is no broader than the evidence supports.
## Progressive Resources
- Deep matrix guide: `references/verification-matrix-and-evidence-strength.md`
- Existing falsification heuristics: `references/falsification-heuristics.md`
- Evidence recording: `references/evidence-recording.md`
- Example: `examples/falsification-session.md`
Referenced files: 8
get-fable8.79 KB
---
name: get-fable
description: "Orchestrate software engineering workflows across the canonical get-fable coding lifecycle with deterministic routing and evidence precedence. Use when starting a complex coding task, navigating lifecycle phases, resuming work with durable .fable state, or routing between research, planning, testing, verification, and recovery — even if the user does not explicitly say \"get-fable\" (e.g. \"follow fable lifecycle\", \"orchestrate this project\", \"what is the next engineering step\", \"route my task\"). Do NOT use when an individual specialist skill already has clear isolated ownership of a bounded subtask."
version: 1.3.0
pack: core
inputs:
- task_description
- current_state
requires:
- repo_access
produces:
- routing_decision
gates:
- state_schema_valid
fallback: null
mutatesWorkspace: false
parallelSafe: true
neural_links:
precursors: []
continuations:
- fable-discover
- fable-research
- fable-plan
lateral_peers:
- fable-spark
recovery: fable-recover
---
# get-fable
Choose the next engineering mode from the actual state of the work, not from the loudest keyword in the user's last message.
## Mission
`get-fable` is the lifecycle orchestrator. It decides **which specialist should own the next decision**, preserves continuity across long sessions, and prevents later phases from skipping evidence that earlier phases have not earned.
Good routing is stateful. "Ship it" does not mean release if verification is stale. "Fix it" does not mean execute if the same hypothesis already failed twice. "Use this API" does not mean code from memory if the contract is current and unknown.
## Activate When
- a substantial engineering request enters the workspace;
- the next Skill is ambiguous;
- work resumes from `.fable` state or a handoff;
- a phase transition is requested;
- new evidence invalidates the current route;
- a failure/review/security/release gate may override ordinary execution.
## Do Not Activate When
- a specialist is already executing a bounded, still-valid contract and no routing condition changed;
- the user asks only for a raw command output that a current Skill already owns;
- repeated re-routing would add ceremony without changing the next safe action.
## Routing Classification
Classify the dominant reason the next action exists.
| Situation | Preferred owner |
| --- | --- |
| repository/runtime facts unknown | `fable-discover` |
| current external fact/API uncertain | `fable-research` |
| architecture/decomposition/contract choice | `fable-plan` |
| testable behavior change | `fable-tdd` |
| genuinely independent bounded cards | `fable-delegate` |
| bounded accepted implementation | `fable-execute` |
| fresh falsification required | `fable-verify` |
| independent diff correctness review | `fable-review` |
| trust boundary/vulnerability/security work | `fable-security` |
| repeated/contradictory failure | `fable-recover` |
| candidate is ready for distribution decision | `fable-release` |
| another session/agent must resume | `fable-handoff` |
| agent-control change needs benchmark proof | `fable-eval` |
| next atomic move is unclear inside active work | `fable-spark` |
## Orchestration Protocol
### Stage 1 — Read intent and durable state
Consider together:
- user's current request;
- active card/phase;
- failure streak;
- mutation vs verified generation;
- open review/security findings;
- unresolved load-bearing unknowns;
- release/handoff state.
The last sentence in chat does not erase the lifecycle state.
### Stage 2 — Apply precedence gates
Before intent scoring, check hard overrides:
1. corrupted/invalid state → diagnose/repair state before trusting it;
2. repeated failure or contradictory evidence → recover;
3. explicit security-sensitive request → security specialist;
4. stale verification when completion/release is requested → verify;
5. blocking unknown that changes design → discover/research;
6. otherwise route by task ownership.
Precedence exists to stop a plausible but unsafe lower-level action.
### Stage 3 — Distinguish task type from requested outcome
Examples:
- "Publish this" is an outcome; current state may still require verify/review first.
- "Fix this" is an outcome; unknown root location may require discovery.
- "Make it faster" may be research/measurement/planning before execution.
- "Review and fix" is two stages; review should identify grounded findings before mutation unless user explicitly asks for direct repair and evidence is already clear.
### Stage 4 — Route to one primary owner
Choose the Skill that owns the **next load-bearing decision**, not every Skill that might eventually participate.
Return:
- selected Skill;
- reason/precedence;
- evidence/state used;
- gate that will permit the next transition.
### Stage 5 — Persist only meaningful transitions
Update lifecycle state when the route changes actual work phase or evidence freshness. Do not churn state for read-only explanatory turns.
### Stage 6 — Re-route on new evidence
A route is not permanent. Recompute when:
- an assumption is disproved;
- scope expands;
- a failure repeats;
- security risk appears;
- mutation makes proof stale;
- a worker discovers a dependency collision;
- release candidate changes.
## Decision Rules
- `failureStreak >= 2` with materially similar attempts outranks execution and routes to recovery.
- Explicit vulnerability/threat/authz/secret/trust-boundary work routes to security even if the diff also needs general review later.
- Current external API/version uncertainty routes to research; repository-local uncertainty routes to discovery.
- If implementation cannot proceed without choosing a new contract/architecture, plan before execute.
- A testable bug/behavior change should route through TDD unless a valid regression harness is impossible or already established.
- "Done", "merge", "ship", or "release" cannot bypass stale/missing verification.
- A passing receipt/research note is not completion-capable evidence.
- Delegation is selected for real independence, not simply because there are multiple subtasks.
- Handoff is selected when continuity itself is the deliverable; it does not substitute for verification.
- Prefer the narrowest specialist that owns the next decision; avoid keeping the orchestrator active once ownership is clear.
## Invariants
- One primary Skill owns the next load-bearing decision.
- State/evidence precedence can override textual intent when necessary for correctness.
- No completion/release route relies on evidence older than relevant mutation.
- Repeated failed hypotheses do not route back into blind execution.
- Research and discovery are distinct: external current fact vs repository/runtime fact.
- Routing explanations remain traceable to intent/state, not hidden scoring alone.
## Failure Taxonomy
### Keyword capture
Router sees "release", "security", or "test" and ignores the actual state/task. Re-evaluate precedence and ownership.
### State blindness
Router uses the last prompt but ignores failure streak, stale proof, or active card. Re-read durable state.
### Over-orchestration
Every small step returns through the router although specialist ownership remains valid. Keep the active specialist until a routing condition changes.
### Under-routing
Execution continues after a new architecture unknown, repeated failure, or security boundary appears. Stop and re-route.
### Multi-owner ambiguity
Several Skills seem plausible. Select the one that owns the earliest unresolved decision; encode later Skills as continuations/gates, not simultaneous primary owners.
### Stale-state corruption
State does not match repository reality. Diagnose/repair before using it as a routing oracle.
## Anti-Patterns
- routing by one keyword;
- treating user's desired final outcome as the immediate next action;
- re-routing every turn for ceremony;
- using Spark as a substitute for specialist selection;
- sending a repeated failure back to execute;
- sending unknown external API details to repository discovery;
- accepting stale verification because the user said "ship it";
- selecting multiple primary Skills with no order.
## Routing Packet
```text
Intent/outcome:
Current phase/card:
State overrides: failure / stale evidence / security / unknowns
Primary next decision:
Selected Skill:
Why this Skill now:
Gate to leave it:
Likely continuation(s):
```
## Completion Criteria
Orchestration for a turn is complete when:
- the next load-bearing decision has one clear owner;
- precedence gates were checked;
- route is consistent with current evidence/state;
- specialist receives enough context to begin without reinterpreting the whole conversation;
- no later lifecycle claim is implied before its evidence exists.
## Progressive Resources
- Deep guide: `references/stateful-routing-and-precedence.md`
- Existing matrix: `references/lifecycle-routing-matrix.md`
- Example: `examples/lifecycle-walkthrough.md`
Referenced files: 8
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- MIT
- Package author
- Mamdouh Abo Ammar
- Keywords
- codex, chatgpt, coding-agent, agent-workflow, coding-lifecycle, verification, skills
Declared capabilities
- Route coding work across 25 canonical lifecycle skills in 8 packs
- Track mutation-aware verification and evidence freshness
- Use specialized TDD, review, security, and recovery specialists
- Run lifecycle hooks across supported Codex and ChatGPT plugin surfaces
- Predict atomic next moves with Fable Spark situational awareness
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 18:00 UTC
- Collection status
- Collected
plugins_6a7da17696b081918e2d9debd654a099
Download plugin data (JSON)