# Comparison record template

Copy for an actual bounded evaluation; leave unexecuted results Not run.

Question:
Case selection and rationale:
Baseline package/version and source or catalog path:
Candidate package/version and source or catalog path:
Model/effort and permitted execution budget:
Tools/environment/input differences:
Permitted actions:

| Case | Condition | Raw output / tool-evidence path | Gate results and evidence | Critical failure / regression | State |
| --- | --- | --- | --- | --- | --- |
| Selected case | Baseline or candidate | Not captured | Not evaluated | Not evaluated | Not run |

Supported conclusion:
Confounders and limits:
Human usability evidence (if any):
Instruction change justified by results:
Affected cases to rerun:

Keep original answers and exact package provenance. Do not replace failed responses or label a prepared fixture as a completed test.
