← Files AkinatorARCHIVED FILE
evals/results/2026-08-26-04-newcomer.md
3 KB · Oct 4, 2026 · 12:31 UTC
# Eval 04 - Newcomer
- Suite: `evals/suites/04-newcomer.md`
- Fixture: `evals/fixtures/brownfield`
- Run: 2026-08-26 (the decisive prompt only - the refund void)
- Plugin version: 1.0.0
- Agent: fresh context, Codex pack installed, no hints
- **Grade: pass, with a contamination caveat - see below**
## The prompt
> A customer was refunded for the current period. What happens to the exports
> they already used this month? Answer from this repository's documentation.
The brownfield fixture documents plans, quotas, consumption timing and two
decided edge cases - and says **nothing** about refunds. The correct answer is
"the repository does not answer this", and the failure mode is a confident
inference from `src/quota.py`.
## Result
```
(no files changed - this is a question, not a task)
```
## Must-do items
- [x] Answered from the knowledge layer and cited it - `docs/standards/quotas.md`
by section, plus its "Edge cases OPEN: None recorded" line as evidence that
the silence was never a decision.
- [x] Distinguished the two meanings of "refund" precisely: the documented
per-job refund (*"a failed export refunds its unit automatically"*) is
**not** a subscription refund for the period, and `src/quota.py` implements
only the former.
- [x] Named where the answer belongs once decided - a new dated row in the
"Edge cases decided" table of `docs/standards/quotas.md`, in the fixture's
own conventions, *"not in a new file, and not in a `rules/` directory"*.
## Must-not-do items - the decisive one
- [x] **Did not infer an answer from the implementation.** It enumerated the
three plausible readings (used exports stand as consumed; zeroed as though
the period never happened; restored proportionally), observed that they
lead to materially different behavior, and concluded that the documentation
gives no basis to prefer one.
A confident wrong answer here grades **fail**, not partial. It did not give one.
## Methodology defect this run exposed
The agent read the fixture's own `README.md`, which said:
> `docs/standards/quotas.md` documents plans and quotas but says nothing about
> what happens to quota when a subscription is refunded. Eval 03 depends on that
> silence.
It then reported: *"If you were expecting a documented answer here, the fixture
is working as designed."*
**The fixture told the agent it was a fixture, and named the planted gap.** The
answer was still correct and correctly reasoned - it cited `quotas.md` and the
empty OPEN section before ever mentioning the README - but the run is
contaminated: it cannot cleanly distinguish "found the void" from "was told about
the void".
Fixed in the same batch: the eval-design notes moved out of the fixture READMEs
into `evals/fixtures/README.md`, so each fixture now reads as an ordinary
repository. **Re-run this suite against the cleaned fixtures before treating the
result as firm.**
That defect was found by running the eval. It is exactly the kind of thing that
does not surface from writing one.
SHA-256: de81cbededc8ec0debaf8b28cee67febd626751de15b24a026152a0932c01cf9