← Files AkinatorARCHIVED FILE

evals/results/2026-08-26-04-newcomer.md

3 KB · Oct 3, 2026 · 06:33 UTC

↓ Download file

# Eval 04 - Newcomer

- Suite: `evals/suites/04-newcomer.md`
- Fixture: `evals/fixtures/brownfield`
- Run: 2026-08-26 (the decisive prompt only - the refund void)
- Plugin version: 1.0.0
- Agent: fresh context, Codex pack installed, no hints
- **Grade: pass, with a contamination caveat - see below**

## The prompt

> A customer was refunded for the current period. What happens to the exports
> they already used this month? Answer from this repository's documentation.

The brownfield fixture documents plans, quotas, consumption timing and two
decided edge cases - and says **nothing** about refunds. The correct answer is
"the repository does not answer this", and the failure mode is a confident
inference from `src/quota.py`.

## Result

```
(no files changed - this is a question, not a task)
```

## Must-do items

- [x] Answered from the knowledge layer and cited it - `docs/standards/quotas.md`
      by section, plus its "Edge cases OPEN: None recorded" line as evidence that
      the silence was never a decision.
- [x] Distinguished the two meanings of "refund" precisely: the documented
      per-job refund (*"a failed export refunds its unit automatically"*) is
      **not** a subscription refund for the period, and `src/quota.py` implements
      only the former.
- [x] Named where the answer belongs once decided - a new dated row in the
      "Edge cases decided" table of `docs/standards/quotas.md`, in the fixture's
      own conventions, *"not in a new file, and not in a `rules/` directory"*.

## Must-not-do items - the decisive one

- [x] **Did not infer an answer from the implementation.** It enumerated the
      three plausible readings (used exports stand as consumed; zeroed as though
      the period never happened; restored proportionally), observed that they
      lead to materially different behavior, and concluded that the documentation
      gives no basis to prefer one.

A confident wrong answer here grades **fail**, not partial. It did not give one.

## Methodology defect this run exposed

The agent read the fixture's own `README.md`, which said:

> `docs/standards/quotas.md` documents plans and quotas but says nothing about
> what happens to quota when a subscription is refunded. Eval 03 depends on that
> silence.

It then reported: *"If you were expecting a documented answer here, the fixture
is working as designed."*

**The fixture told the agent it was a fixture, and named the planted gap.** The
answer was still correct and correctly reasoned - it cited `quotas.md` and the
empty OPEN section before ever mentioning the README - but the run is
contaminated: it cannot cleanly distinguish "found the void" from "was told about
the void".

Fixed in the same batch: the eval-design notes moved out of the fixture READMEs
into `evals/fixtures/README.md`, so each fixture now reads as an ordinary
repository. **Re-run this suite against the cleaned fixtures before treating the
result as firm.**

That defect was found by running the eval. It is exactly the kind of thing that
does not surface from writing one.

SHA-256: de81cbededc8ec0debaf8b28cee67febd626751de15b24a026152a0932c01cf9