← get-fableCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to get-fable
Snapshot Sep 30, 2026 · 23:14 UTC · version 1.5.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Evaluate changes to agent prompts, skills, routing policies, and harnesses against reproducible baselines, held-out suites, and regression benchmarks. Use when optimizing agent system prompts, measuring skill triggering accuracy, evaluating routing changes, or running benchmark regressions — even if the user does not explicitly say \"fable-eval\" (e.g. \"benchmark this prompt\", \"evaluate agent behavior\", \"test skill performance\", \"run the eval suite\"). Do NOT use for routine application unit tests (use fable-verify).",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 387
},
{
"relative_path": "evals/scenarios.json",
"size_in_bytes": 3967
},
{
"relative_path": "examples/eval-regression-run.md",
"size_in_bytes": 276
},
{
"relative_path": "references/benchmark-design-and-overfit-control.md",
"size_in_bytes": 2722
},
{
"relative_path": "references/eval-harness-protocol.md",
"size_in_bytes": 1656
},
{
"relative_path": "skill.package.json",
"size_in_bytes": 461
},
{
"relative_path": "templates/eval-benchmark-report.template.md",
"size_in_bytes": 835
}
],
"name": "fable-eval",
"skill_md_contents": "---\nname: fable-eval\ndescription: \"Evaluate changes to agent prompts, skills, routing policies, and harnesses against reproducible baselines, held-out suites, and regression benchmarks. Use when optimizing agent system prompts, measuring skill triggering accuracy, evaluating routing changes, or running benchmark regressions — even if the user does not explicitly say \\\"fable-eval\\\" (e.g. \\\"benchmark this prompt\\\", \\\"evaluate agent behavior\\\", \\\"test skill performance\\\", \\\"run the eval suite\\\"). Do NOT use for routine application unit tests (use fable-verify).\"\nversion: 1.3.0\npack: evolution\ninputs:\n - candidate_modification\nrequires:\n - reproducible_baseline\nproduces:\n - eval_verdict\n - regression_evidence\ngates:\n - baseline_frozen\n - holdout_tested\n - rollback_defined\nfallback: fable-plan\nmutatesWorkspace: false\nparallelSafe: true\nneural_links:\n precursors:\n - skill-creator\n continuations:\n - fable-plan\n - fable-execute\n lateral_peers:\n - fable-verify\n recovery: fable-recover\n---\n\n# Fable Eval\n\nMeasure whether an agent-control change improves the behavior it claims to improve without quietly overfitting the benchmark or breaking neighboring behavior.\n\n## Mission\nAn eval is a decision instrument, not a scoreboard. It needs a frozen comparison point, representative semantic families, oracle isolation, explicit failure costs, and a rollback decision.\n\nA candidate should not win because the prompts resemble its instructions, because holdouts leaked into authoring, or because one average score hides a severe regression.\n\n## Activate When\n- changing Skills, prompts, routers, hooks, agent profiles, policies, or model-control logic;\n- comparing candidate prompt/agent configurations;\n- measuring trigger precision/recall or action compliance;\n- validating a new behavioral maturity claim;\n- investigating whether an apparent improvement is robust or benchmark-specific.\n\n## Do Not Activate When\n- verifying ordinary application behavior (`fable-verify`);\n- authoring a Skill before its intended behavior is clear (`skill-creator`);\n- running a one-off subjective prompt demo with no acceptance decision.\n\n## Evaluation Classification\n| Change | Primary eval risk |\n| --- | --- |\n| Router/trigger | false positives, false negatives, precedence |\n| Skill instruction | action correctness, forbidden shortcuts, boundary behavior |\n| Spark/next-action | top-1 action, unsafe suggestion, silence precision |\n| Hook/guard | enforcement, false blocking, bypasses |\n| Prompt/persona | task quality + regressions + instruction conflicts |\n| Tool policy | correct tool choice, unsafe/missing action |\n| Model/config | quality/latency/cost variance across representative tasks |\n\n## Protocol\n### Stage 1 — Define the decision before running tests\nState:\n- candidate being evaluated;\n- baseline/control;\n- exact behavior expected to improve;\n- metrics and thresholds;\n- unacceptable regressions;\n- rollback action.\n\nAvoid inventing metrics after seeing results.\n\n### Stage 2 — Build semantic scenario families\nCover distinct decisions, not wording variants. Include as applicable:\n- straightforward positive case;\n- non-trigger/boundary case;\n- ambiguous competing action;\n- adversarial shortcut pressure;\n- partial/contradictory evidence;\n- failure/recovery path;\n- legacy/constrained environment;\n- unseen holdout.\n\nRecord family coverage separately from raw prompt count.\n\n### Stage 3 — Freeze baseline and corpus identity\nBind the run to:\n- corpus hash/version;\n- candidate/baseline identity;\n- evaluator/scorer version;\n- provider/model/config where external;\n- timestamp/repository revision.\n\nChanging the subject, oracle, corpus, or scoring logic invalidates direct comparability unless explicitly normalized.\n\n### Stage 4 — Protect the oracle\nProvider-facing requests must not reveal expected actions, forbidden actions, category labels, holdout identity, scoring implementation, or answer-bearing metadata.\n\nDo not draft the candidate while repeatedly reading holdout failures. Promote discovered cases into a future checked corpus and preserve a new unseen holdout.\n\n### Stage 5 — Execute baseline and candidate consistently\nUse the same task inputs, tool availability, context budget, temperature/configuration, timeout policy, and scoring rules where comparison requires them.\n\nCapture provider errors/timeouts as failures or explicit unavailable states; do not fill missing outputs from the oracle.\n\n### Stage 6 — Score by slices, not average alone\nInspect:\n- overall metric;\n- each semantic family;\n- negative/adversarial forbidden violations;\n- high-cost regressions;\n- variance/repeated-run stability where stochasticity is material;\n- routing confusion pairs where applicable.\n\nA 1% average gain is not acceptable if it introduces a severe release/security/recovery regression.\n\n### Stage 7 — Investigate suspicious gains\nCheck for:\n- prompt leakage;\n- duplicated/near-duplicate scenarios;\n- benchmark-specific keyword matching;\n- changed tool/context budget;\n- scorer drift;\n- cherry-picked seeds/runs;\n- examples copied into the candidate.\n\n### Stage 8 — Decide and preserve rollback\nVerdict:\n- **ACCEPT**: thresholds met, no prohibited regression, evidence representative/fresh;\n- **REJECT**: candidate regresses required behavior or fails threshold;\n- **INCONCLUSIVE**: evidence lacks breadth/stability/holdout integrity.\n\nRecord baseline artifact so rollback remains possible.\n\n## Decision Rules\n- Semantic family breadth matters more than raw scenario count.\n- Surface rewrites of one case do not create independent coverage.\n- Holdouts stop being holdouts once used repeatedly to tune the candidate.\n- Compare slices before averages; safety-critical forbidden violations can veto a higher average score.\n- If stochastic variance could change the decision, repeat enough runs to estimate stability rather than cherry-picking one seed.\n- If provider/runtime errors differ between candidate and baseline, separate infrastructure failure from behavior score.\n- Never preserve an old maturity result after the evaluated Skill/corpus/oracle changes unless freshness validation proves identity.\n- Do not lower thresholds after a candidate fails simply to ship it.\n\n## Invariants\n- Baseline is reproducible/frozen before candidate judgment.\n- Oracle/holdout data remains hidden from the evaluated agent.\n- Candidate and baseline are compared under equivalent conditions where claimed.\n- Missing/failed provider outputs are never replaced by expected answers.\n- Every acceptance decision has a rollback path.\n- High-cost regressions remain visible even when aggregate score improves.\n\n## Failure Taxonomy\n### Benchmark overfit\nCandidate improves checked cases but fails unseen family/holdout. Increase semantic breadth and reject/generalize candidate.\n\n### Oracle leakage\nExpected/forbidden/category data reaches provider/candidate authoring loop. Discard contaminated evidence and create fresh blind cases.\n\n### Metric blindness\nAverage improves while important slice worsens. Use per-family and veto metrics.\n\n### Non-comparable runs\nDifferent model/tool/context/scorer settings produce apparent gain. Re-run under controlled conditions.\n\n### High variance\nRepeated runs change verdict. Increase samples/control nondeterminism or mark inconclusive.\n\n### Corpus drift\nSkill/scenario/oracle changed after evidence capture. Mark evidence stale and rerun.\n\n## Anti-Patterns\n- five paraphrases counted as five independent tests;\n- reading holdouts while tuning every candidate;\n- accepting on average score alone;\n- changing thresholds after seeing failure;\n- treating provider timeout as skipped rather than failed/incomplete;\n- evaluating only positive examples;\n- copying expected action vocabulary in a way that reveals the answer per case;\n- claiming M4 because an evidence JSON file exists even though corpus hash changed.\n\n## Eval Report\n```text\nCandidate / baseline:\nCorpus + hashes:\nProvider/config:\nSemantic families:\nMetrics + thresholds:\nPer-family results:\nForbidden/high-cost regressions:\nVariance/repeats:\nLeakage/comparability checks:\nVerdict: ACCEPT | REJECT | INCONCLUSIVE\nRollback:\nEvidence freshness:\n```\n\n## Completion Criteria\nEvaluation completes when:\n- baseline and candidate identities are explicit;\n- semantic coverage and blind holdout integrity are credible;\n- metrics are inspected by meaningful slices;\n- regressions/forbidden actions are not hidden by averages;\n- verdict and rollback are evidence-backed;\n- evidence freshness is tied to the exact evaluated corpus/control.\n\n## Progressive Resources\n- Deep guide: `references/benchmark-design-and-overfit-control.md`\n- Existing protocol: `references/eval-harness-protocol.md`\n- Example: `examples/eval-regression-run.md`\n"
}SHA-256 of public snapshot: d677dbb0efe7419e283d8bfb975755be6f0486c60b09a43de38f857b95a974c0