# Oracle Independence and Simulation Boundaries

An oracle is useful only to the extent that it can be wrong in a different way from the candidate.

## Independence checklist
Ask:
- Was expected behavior derived from a spec/public contract rather than candidate internals?
- Does reference code use a different algorithm/implementation path?
- Are golden fixtures sourced from trusted historical/runtime observations?
- Does a secondary model/provider receive only legitimate task context, not expected answers?
- Which assumptions are shared between oracle and candidate?

If candidate and oracle share the same parser, mock, fixture generator, or mistaken interpretation, agreement can be false confidence.

## Oracle types

### Golden fixtures
Strong when provenance is trustworthy and the contract is stable. Weak when fixtures were generated by the current implementation or silently updated after failures.

### Reference implementation
Prefer small, obviously correct, separately derived logic. A second copy of complex production code is not much of an oracle.

### Property/invariant oracle
Useful for algebraic/structural rules such as round trips, ordering, conservation, idempotency, monotonicity.

### External/reference provider
Useful for differential behavior if requests are blinded and provider/model identity is recorded. Agreement is still evidence, not mathematical proof.

### State-machine model
Useful for workflows. Define allowed transitions and invariants from product contract, then compare candidate transitions and side effects.

## Differential mismatch handling
Do not decide by majority or by assuming the reference is right. Re-ground the disputed behavior in the authoritative contract/source. Sometimes the mismatch discovers a stale fixture or reference bug.

## Simulation boundary
Always label:
- what is simulated/faked;
- what is real code/runtime;
- what external behavior is assumed;
- which faults were injected;
- which real-world evidence remains required.

For example, a fake payment gateway can prove adapter handling of a documented timeout contract, but not that the real gateway currently emits the same edge response in your account/region.

## UI causal evidence
Prefer interaction evidence:
`click → request/state transition → visible result`
plus console/network errors where relevant. Pixel comparison alone is useful for visual regressions, not complete functional correctness.

## Workspace safety
Simulations should use temporary directories, fixture databases, disposable ports, and sandbox state. Never remove untracked user files to obtain a clean test environment.