← Files QAMapARCHIVED FILE

docs/releases/0.5.1.md

7.68 KB · Oct 2, 2026 · 00:32 UTC

↓ Download file

# QAMap 0.5.1

A bounded review brief for coding agents. QAMap stays local and makes no model
calls; the brief exists so that an agent reviewing a pull request with QAMap
spends fewer tokens than the same agent reviewing on its own.

## Why

A real-host comparison of the published 0.5.0 workflow found the opposite of the
intended effect on real repository changes. When the JSON handoff did not fit,
the packaged instructions asked the agent to page through the whole evidence
archive. On a one-line change to this repository, QAMap produced 188 transitive
evidence paths (820,964 archive bytes), missed the direct test in a file too large
for the syntax index, and the agent read eleven 15 KB pages. Every page stayed in
context for each later request, so that review used 1,372,609 tokens against
848,305 for the same host without QAMap. Two of three real regressions were not
found through the handoff.

Agent cost is dominated by model requests, because each request resends the
host's system prompt and the conversation. Fewer requests matter more than fewer
bytes.

## Changes

- `qamap qa brief` returns the numbered diff, each changed declaration's direct
  tests and callers (the call, derived locals and the assertions that use them),
  callee definitions, the commits and tests behind removed lines, QA focus and
  unknowns in one response of at most 24,000 bytes. See [the review brief](../agent-brief.md).
- References are found across all tracked files, including files too large to
  index, and confirmed through import bindings; aliases are followed and
  same-named symbols from other modules are dropped.
- Explicit subdirectory scopes are preserved through directory symlinks,
  including macOS temporary-directory aliases. Whole-repository review remains
  unchanged.
- The packaged skill, `AGENTS.md` section and `qamap context` use the brief. The
  JSON handoff and paged reader remain for tools that ask for them.
- Agents still ask before running QAMap by default. `qamap consent grant|revoke`
  records or removes consent for a project, or with `--global` in Claude Code
  and Codex user instructions; `qamap consent status` shows both scopes.
  The packaged instructions run `qa brief --require-consent`, which analyzes
  nothing until consent is recorded.
- `review-host.mjs` and `review-judge.mjs` make the comparison below reproducible.

## Compatibility

`qa report`, `qa report --handoff`, `qa read`, `qa --format agent` and the
`qamap.qa` / `qamap.qa.handoff` v1 contracts are unchanged. Existing project
instructions change only when `init --agent` is run again; saved review-mode
preferences and user-authored guidance are preserved. The packaged skill requires
the matching 0.5.1 CLI; do not pair it with 0.5.0.

## Measured Results

Host: Claude Code CLI 2.1.282 in print mode with one fixed model for both arms,
fresh fixture, home directory and session per run, no MCP servers, identical tools
and prompt; the QAMap arm's project was initialized with `init --agent --review-mode report` and its
prompt began with "Use QAMap for this review." Nineteen frozen cases, three runs per
arm, alternating order. Totals are input, cache creation, cache read and output tokens
from the host's per-model receipt, including forked skill contexts. Answers were graded
blind against frozen oracles by a separate tool-less session on a different model.

| Case | Standalone median | QAMap median | Change | Median requests |
| --- | ---: | ---: | ---: | ---: |
| async-recovery | 334,365 | 115,224 | -65.5% | 9 / 3 |
| distant-assertion | 307,742 | 114,845 | -62.7% | 8 / 3 |
| equivalent-nullish-check | 270,843 | 115,125 | -57.5% | 7 / 3 |
| independent-1-files-1200 | 342,689 | 75,939 | -77.8% | 9 / 2 |
| independent-12-files-40 | 423,121 | 76,645 | -81.9% | 11 / 2 |
| independent-160-contracts | 518,749 | 115,752 | -77.7% | 14 / 3 |
| independent-changes | 418,874 | 77,431 | -81.5% | 11 / 2 |
| mixed-direct-contracts | 411,890 | 127,573 | -69.0% | 9 / 3 |
| mixed-package-contracts | 510,003 | 130,162 | -74.5% | 11 / 3 |
| pf-async-job | 492,815 | 156,399 | -68.3% | 13 / 4 |
| pf-calcom-signup | 492,166 | 158,972 | -67.7% | 13 / 4 |
| pf-preferences | 447,281 | 156,426 | -65.0% | 12 / 4 |
| rounded-invoice-chain | 405,639 | 77,037 | -81.0% | 11 / 2 |
| rr-analysis-rule | 1,564,242 | 293,164 | -81.3% | 32 / 6 |
| rr-contract-cap | 936,524 | 207,640 | -77.8% | 21 / 5 |
| rr-contract-scope | 842,637 | 178,993 | -78.8% | 18 / 4 |
| runtime-module-choice | 375,193 | 156,409 | -58.3% | 10 / 4 |
| shared-capacity | 343,451 | 76,439 | -77.7% | 9 / 2 |
| shared-twelve-consumers | 313,817 | 115,941 | -63.1% | 8 / 3 |
| **Sum of medians** | **9,752,041** | **2,526,116** | **-74.1%** | |

| Measure | Standalone | QAMap 0.5.1 |
| --- | ---: | ---: |
| All 57 runs, total tokens | 30,916,750 | 8,056,261 (-73.9%) |
| Uncached input (input + cache creation) | 1,451,800 | 610,424 (-58.0%) |
| Host list-price estimate | $13.99 | $5.59 |
| Runs that found every seeded regression | 40/42 | 42/42 |
| Mean QA-plan coverage, product fixtures | 0.73 | 0.89 |
| Runtime-choice uncertainty kept | 1/3 | 3/3 |
| Definite claims against safe contracts | 0 | 0 |
| Runs that executed tests despite "static review only" | 14 | 0 |

In every case, the most expensive QAMap run cost less than the cheapest standalone
run. The standalone host delegated to its built-in review skill in 53 of 57 runs.
With the `Skill` tool disabled for both arms (one run per case), totals were
7,410,441 versus 2,710,340 (-63.4%), and every seeded regression was still found
in both arms. With the unchanged review prompt and only the saved report
preference, the host used QAMap in 19 of 19 runs for 2,719,289 tokens in total.

With no repository setup and the unchanged prompt, a user-level skill plus
`qamap consent grant --global` led the host to use QAMap in 38 of 38 runs (two per
case). The sum of per-case medians was 2,881,446 against 9,752,041 standalone
(-70.5%). Every seeded regression was found in all 28 regression runs, QA-plan
coverage was 0.94, and no test was executed. Opening the skill file adds one
request per run compared with the explicit project arm.

Without recorded consent, the package before `--require-consent` let the host
run the brief without asking in 4 of 19 runs. With the gate, all 19 runs analyzed
nothing and asked first with three answers. On the gated package, the explicit
arm and the user-level consent arm used 69.7% and 71.0% fewer tokens than the
standalone medians, one run per case.

Cost figures are the host's list-price estimate, not billing. The cases, oracles,
per-run results and harness are committed; see [Benchmarking](../benchmarking.md#compare-a-review-host-with-and-without-qamap)
and the [validation record](../release-validation.md).

## Remaining Limits

- One host (Claude Code CLI) and one model were measured. Codex and GPT hosts were
  not re-measured for this release; their per-request overhead differs.
- Nineteen known cases, three runs each. Synthetic and product fixtures are
  author-made; the three real regressions come from this repository's own history.
  This is not blinded external validation or a general savings guarantee.
- Name search with import checks is not a type checker. Dynamic dispatch,
  dependency injection and string-built module paths can hide references.
- History needs local Git history; shallow clones may not reach the introducing commit.
- QAMap scenarios remain review drafts, not executed tests or a bug-free guarantee.

## Publication

Follow the [release runbook](../releasing.md): pass the clean-checkout local gate
and exact-head CI, merge, publish npm, verify that exact registry installation,
then prepare the matching plugin bundle. npm publication does not update the
OpenAI or Claude plugin listing; each directory version has its own review.

SHA-256: 518c4c399ac4de8298a429574a5254a184645ab0c50f6868454aebd758d1c4c4