← Files Compound EngineeringARCHIVED FILE
docs/solutions/skill-design/authored-eval-corpora-contain-the-happy-path.md
4.43 KB · Oct 4, 2026 · 12:33 UTC
--- title: "An authored eval corpus contains the mechanism's happy path" date: 2026-08-13 category: skill-design module: skill-evaluation problem_type: workflow_issue component: evaluation severity: high tags: - skill-eval - evaluation-design - corpus - blind-scoring - false-pass --- # An authored eval corpus contains the mechanism's happy path ## What happened A reasoning-based duplicate matcher for `ce-doc-review` was evaluated against a purpose-built corpus: plans generated by an independent harness, with defect manifests withheld from the session that wrote the matcher, scored by an independent grader. The blinding was real and carefully maintained. It returned **100% merge recall and 100% merge precision** across every host and model cell. A later run measured the same matcher — unchanged bytes — against a duplicate pair captured from a real review. **Merge recall was 20-36%**, mean 26%. One trial also produced two false merges. The matcher had not regressed. The corpus was easy. ## Why it happened The corpus was generated with an instruction to plant cross-persona duplicate pairs. A duplicate authored *so that a matcher can find it* is a duplicate whose two halves describe the same problem in recognisably parallel terms. Organic duplicates are not like that: two reviewers arrive at one problem from different lenses, attach it to different sections, describe it in vocabulary drawn from their own briefs, and propose fixes that overlap without matching. That is the case the matcher exists for, and it is the case the corpus did not contain. The blinding protected against the wrong failure. Withholding the manifest stops the **scorer** from being biased by the mechanism's author. It says nothing about whether the **inputs** are representative. A corpus can be perfectly blind and still be a soft target. ## The cost Four evaluation rounds ran against that corpus before anyone tested real material. Worse, a downstream design decision — cancelling a decision-clustering stage — rested on a decision-load number read off the same inputs. If duplicates survive merging at 26% rather than 0%, the post-merge load is far higher than measured and the cancellation reasoning does not hold. One unrepresentative corpus invalidated a measurement and a design decision built on it. ## Sampling real artifacts has its own version of this The obvious correction — stop generating, sample real documents — was also applied in the same effort, and it carried a different bias in the same direction. The plans sampled were **already implemented and had already been through review**. A document that has been reviewed has had its defects removed, so it under-produces findings and under-produces duplicates; the measured duplication rate was a floor, not the rate. So "use real artifacts" is not sufficient on its own. What you want is artifacts in the state the skill actually meets them — for a review skill, documents that have **not** yet been reviewed. When only post-review material is available, the *causal* results still hold (which mechanism failed, and why) while the *rate* results do not. Say which of the two you are relying on. ## What to do instead - **Prefer captured real artifacts over generated ones, in the state the skill meets them.** Agent sessions leave transcripts; real reviews, real findings, and real user reactions can be extracted from them. A recorded run where a human pushed back is worth more than a generated fixture, because the ground truth was produced by someone with no stake in the mechanism. Check the artifact has not already been cleaned by the process you are testing. - **When you must generate, do not let the mechanism's author set the difficulty.** Have the generator choose the defect mix and counts itself, and include cases where the mechanism should find nothing. - **Validate the corpus against at least one real case before trusting anything measured on it.** A single organic example is enough to catch a soft target — it was here. Run that check first, not fourth. - **Treat a perfect score as a corpus smell.** 100% across every cell is more often evidence about the inputs than about the mechanism. ## Related - A frozen finding set cannot measure an emission-layer change: see Leak C in [`paired-old-vs-new-injection-skill-evals.md`](paired-old-vs-new-injection-skill-evals.md). - Trial-count and variance discipline: `paired-old-vs-new-injection-skill-evals.md` (technique 8), `ce-doc-review-calibration-patterns.md`.
SHA-256: 7496bf2451f61e5cfbfbed658b9293c3f7f003860e9808b57d195ba7d5751e18