← Files taskplaneARCHIVED FILE

CHANGELOG.md

131 KB · Oct 2, 2026 · 00:29 UTC

↓ Download file

# taskplane changelog

## v2.31.5 — Cowork workspace binding and recovery

- Separate durable workspace storage from execution and worker locality; require explicit Cowork bindings and refuse missing or conflicting roots before creating state.
- Preserve validated same-binding recovery history through retirement and relocation without transferring approval.
- Return structured context errors when the workspace is omitted or unresolved, and reject malformed unavailable-worker evidence before changing grants or state.
- Keep Claude-shaped hooks, native worker identity and workspace contracts aligned; document setup, recovery and manual live-host verification limits.
- Verified with 450 affected regression tests, two host-shape checks, independent Evaluation and parallel Engineering reviews. Live Claude Cowork verification remains a manual follow-up.
- Exercise packaged Windows hook launchers and explicitly test unsupported binding operations without mutation; successful binding fixtures require the actual filesystem primitives.

## v2.31.4 — native workers and bounded context

- Scope native read inputs and freshness per task, with conservative legacy fallback.
- Preflight required context, drain bounded bodies, require explicit history isolation and observed startup before cohort release.
- Expose retries and context cost; accept clear end-to-end automatic-approval wording.
- Bound consumed-root receipt metadata for large contexts and retain full receipt validation.
- Reject future human-approval conditions while allowing required-check conditions.

- Require typed native execution in new-run guidance, with one distinct worker per
  selected review lens and explicit serial exceptions. Seal frozen requirements
  against fresh joined results instead of accepting reviewer labels.
- Preserve exact workspace/run/task context continuations and current locked
  lifecycle state during delivery.
- Route independently bound forks to their own runs while retaining actual worker
  parent authority; fingerprint root verification inputs like native inputs.
- Keep unavailable capacity and legacy unverified coverage explicit. Source and
  package fixtures do not certify an actually loaded host runtime.

## v2.31.3 — scoped native worker restoration

- Add native task reservations, observed identity binding, worker context receipts,
  lifecycle joins and verified root result acceptance under the existing harness.
- Admit ready independent work up to observed capacity, without a default two-worker cap.
- Publish late task definitions through an explicit run-bound context update.
- Record denied/error hook outcomes and early returns without weakening refusals.
- Account for reused guardian/worker intervals and retain missing/reset coverage.
- Expose worker attempts, scheduling reasons and hook outcomes in the shared dashboard.
  Package checks and contract fixtures remain distinct from actual loaded-runtime
  activation and live parallel evidence.

## v2.31.2 — usage remediation, security and workflow recovery

- Preserve inactive workflow history with bounded immutable retention and explicit retirement without acceptance.
- Size context batches against prospective receipts, prevent oversized consume envelopes, and introduce compatible semantic handoffs for new runs.
- Discover archived native sessions and verified fork lineage with scoped error diagnostics.
- Recognize explicit end-to-end auto-approval wording; expose policy bindings in compact results and expand safe diagnostic grammar.
- Preserve completed-run follow-up access when installed execution skills are reread; add explicit native deactivation for uninitialized selections without changing approvals or active-run guards.
- Bound complete serialized harness updates before atomic replacement, preserving valid state on Unicode or combined-field overflow.
- Reject indirect package symlinks and mismatched versions, disable persisted CI checkout credentials, and add repeatable security checks.
- Give these fixes a distinct version in both plugin manifests and marketplace metadata so the marketplace update can identify the new package.

## v2.31.1 — context performance and release identity

- Package the merged context batching and per-run task isolation fixes under a distinct version for reliable marketplace update resolution and tracking.
- Include the merged Windows portability and partitioned CI improvements. Runtime and CI code are unchanged from the verified main build.

## v2.31.0 — explained token accounting and conversational approvals

- Reconcile measured tokens into phase activity, activity outside phases, or explicitly unresolved intervals, with session, boundary and reason details in reports and the dashboard.
- Recover skipped observations within the same phase visit without inventing a work/review split. Keep missing cross-phase boundaries, partial counters and counter resets visible as unresolved.
- Separate pre-run usage and post-completion follow-up, deduplicate observations, and preserve legacy usage fields. A recorded historical replay attributes 3,756,746 formerly unallocated tokens to Build; this is improved attribution, not token savings.
- Accept clear ordinary-language phase approvals and explicit automatic-workflow instructions, including common approval typos. Ask natural clarification for questions, conditions, conflicting instructions or unclear intent without demanding an exact phrase.
- Preserve checkpoint, phase, provenance and presentation binding. Mixed responses such as “Approved. Needs changes.” cannot silently approve a phase.
- Cover accounting conservation, historical boundaries, approval ambiguity, native workflow integration and the dashboard with regression checks.

## v2.30.0 — bounded context and safe verification reuse

- Return bounded structured summaries for routine workflow commands, with explicit full-output access for clients that need the complete state.
- Persist canonical context and phase handoffs with scoped views, verified references, explicit budgets and deduplicated shared evidence.
- Reuse passing checks only when their relevant source, command, environment and runtime evidence still matches. Treat unverified external dependencies as ineligible and preserve independent phase approvals.
- Wire the context interface into delivery and review instructions and both packaged host workflows.
- Migrate legacy recovery test clients to explicit full output so their recovery journeys remain compatible with the compact default.
- Verify a matched seven-phase delivery fixture with 79.8% fewer context bytes across three paired repetitions. This is a deterministic byte measurement, not a claim about live model-token or cost savings; the distinct original canonical-input benchmark remains below its required reduction target.

## v2.29.0 — harness lifecycle repairs

- Keep current Changes requested and Rejected checkpoints editable across multiple corrections and resubmission, while preserving accepted evidence and packet history.
- Preserve observed command identity across approval revisions. Allow same-visit polling, interruption and terminal observations without allowing new code on stale grants or reopening terminal handles.
- Resolve transcript ancestry and observed child-start links before the first child scope check.
- Route bare, plugin-tagged and direct Taskplane requests by their leading action, so later design or review wording cannot override Build.
- Exercise these repairs and negative boundaries against source and both extracted plugin packages, including correction through declared native hooks.

## v2.28.0 — product guidance and reliable workflow recovery

- Restore the product purpose, animated overview and practical design/build/review/status examples. Keep installation and hook trust concise, and include the animation and editable source in both packages.
- Accept bounded inline autonomous assessments so a submitted checkpoint does not require an otherwise prohibited file write.
- Require an exact native dashboard handoff before automatic approval, advance or finish, including standalone phases. Preserve the handoff across an approval-only revision change and allow the exact native opener.
- Keep metadata-only diagnosis, exact Taskplane administration and observed clean-checkout recovery available without loosening implementation scope or changing user hook settings.
- Record per-member package hashes and whether the packaged bytes match the named source commit. A version label or dirty working tree is no longer treated as proof of installed parity.
- Close operation-path test gaps: invoke declared PreToolUse/PostToolUse commands around phase journeys, verify continuous transitions without relying on Stop, and distinguish extracted-package results from live desktop enforcement.

- Start a new explicitly requested run with `flow start --replace-run` even when the old checkpoint is sealed or stale. Preserve its evidence, revoke its grants, and select a fresh native dashboard without carrying over approvals.
- Record exactly bound human changes/rejection/cancellation despite source drift; continue to reject approval or advancement over changed evidence.
- Keep bounded diagnostics, CLI help and fresh recovery-scope preparation available at checkpoints. Reject shell chaining, implementation writes, stale replacement requests and replacement while known processes are running.
- Cover recovery at all seven delivery checkpoints, standalone entries, declared Codex hooks and the extracted Codex package. Document the root cause and missing regression coverage.

## v2.27.0 — onboarding, autonomous approvals and execution harness

- Keep the autonomous build menu prompt within Codex’s supported length so it is shown after installation.
- Keep Codex hook token counters on the Codex transcript reader when the host also exports `CLAUDE_PLUGIN_ROOT` for compatibility.
- Accept Codex native `apply_patch` hook payloads (`tool_input.command`) while retaining bootstrap, phase-scope and sealed-checkpoint enforcement. The installed-host test exposed a contract mismatch hidden by synthetic fixtures.

- Document current Codex/Claude installation, first-task onboarding and post-installation hook review/trust, replacing retired setup commands.
- Keep manual phase approval as the default; add explicit run-bound automatic approval policies, evidence-backed conditions, separate policy decisions, revocation and replay protection.
- Publish native dashboard snapshots for an exact run, with phase/visit and cumulative token coverage, static-refresh instructions, captured graph/task context and workspace-bound graph provenance.
- Validate manual/autonomous native journeys, negative authorization cases, snapshot isolation, counter gaps and extracted Codex/Claude packages.

- Activate the harness before initialization for shipped execution prompts, skills and standalone reviews; enforce current phase evidence and native dashboard handoff on completion and resume.
- Preserve complete checkpoint history with compact storage under the existing limit; distinguish saved counter times, task observations and current dashboard publication.
- Include the onboarding guide in both distributable packages.
- Keep the sealed-checkpoint dashboard current when reviews contain multiple evidence paths; permit the native dashboard handoff with quoted punctuation while continuing to reject shell chaining and arbitrary output paths.
- Exercise the declared Windows launcher during harness initialization and resume; normalize bootstrap paths and keep POSIX permission assertions separate from Windows storage checks.

## v2.26.0 — native workflow gates and shared phase evidence

- Restore explicit human acceptance at every Product, Design, Plan, Build,
  Evaluate, Engineering and Retro checkpoint. Bind observed decisions to the
  presented evidence, source revision, scope and task prerequisites.
- Use the shared dependency graph, source component decomposition, task DAG and
  dashboard in every phase, including standalone Product, Design and Engineering.
  Preserve findings when Engineering starts a human-approved Product route.
- Ship cooperative native workflow controls for Codex and Claude, with explicit
  limits on host authentication, protected storage, tool containment and process
  tracking. Refuse protected-host execution when its required controls are absent.
- Harden journal paths and writes, invalidate stale dependency caches, preserve
  resumed usage baselines and handle literal Git filenames safely. Clarify the
  local data collected and the limits of the privacy boundary.
- Validate workflow state against its actual serialized size, preserve private
  atomic writes, and keep finishing an older run from clearing a newer active run.
  Strengthen the regression check for temporary-file cleanup after failed sync.
- Fetch fresh dashboard fixture pages in browser tests so rapid approval updates
  cannot reuse an earlier HTTP cache response.
- Keep generated lens files byte-identical across platforms and make portability
  tests respect the POSIX-only sandbox profile, native Windows directory flushes,
  and the distinction between POSIX mode bits and Windows ACLs.

## v2.25.0 — shared delivery flow and native telemetry

- Keep the orchestrator responsible for delivery through Product, Design, Plan,
  Build, Evaluate, Engineering and Retro. Telemetry advises without blocking
  tools or imposing a Taskplane token limit.
- Give every stage the same task decomposition, dependency graph, attached
  evidence and dashboard. Refresh the dashboard at meaningful progress updates.
- Render styled dependency columns, module selection, component details and
  change impact alongside execution and review evidence.
- Reconcile Codex root and lens counters across native sessions, including final
  responses. Show missing coverage and host approval usage explicitly.
- Support Claude session binding, streamed-message deduplication, cache accounting
  and subagent transcripts in the same flow. Prefer the loaded Claude runtime
  over stale workspace launchers and validate both host packages.
- Preserve local test reports outside distributed source and align package
  regression checks with the advisory delivery path.

The untagged 2.23.1–2.23.10 and 2.24.0 source candidates are superseded by
this version. Existing tags and historical records remain unchanged.

### Included recovery from the 2.23.1 stateless baseline

- Restore the post-refactor 2.23.1 source baseline, preserving the active v4
  phase runtime and intentional removal of the older runtime and obsolete tests.
- Retain ordinary native source review without automatic delivery setup,
  contracts, graph scans or dashboard obligations.
- Add scoped standalone Product entry with usable document tools, planned
  requirement links, artifact delivery, human input and exact normal finish.
  Implementation writes, sibling/worker release and active-delivery release
  remain refused. Native Codex reads use its existing read-only sandbox.
- Resolve fresh native Codex sessions from their exact transcript identity when
  their session-index entry has not yet been written. Reject abbreviated or
  option-like literal arguments that could change Product control authority.
- Show criterion-level Product DoR, Product DoD and implementation acceptance
  separately. Blank statements fail readiness; scores and approvals do not
  claim that required review or implementation evidence exists.
- Probe quality tools in the same isolated environment used for execution.
  Missing dependencies are reported before evaluation begins.

### Included project-local execution and inline setup

- Human-approved Product and Design amendments can be recorded during delivery,
  including Build and Review, with superseded scope, work and evidence retained.
  Revised Design requires fresh human approval before downstream work resumes.
- Dashboard startup no longer calls review-only filters. Review controls are
  initialized only on findings pages; executable regression checks cover
  Dashboard navigation, approval delivery, review filtering, and paged gates.
- New execution defaults to the project's ignored `.taskplane/`, including
  private knowledge, hook receipts, run state, evidence, and managed workspaces.
  Existing bindings remain authoritative; explicit project selection preserves
  unused preflight history and refuses to reset active runs.
- Codex setup uses plugin-provided hooks and restores only the ignored local
  launcher. It no longer registers a duplicate project hook set or requires
  that duplicate for readiness.
- Onboarding uses a Dashboard-style inline form for storage, knowledge sharing,
  and project context. Submissions carry actual values and stale-write checks;
  progress changes only when the engine returns verified readiness. Common
  model and reasoning controls save supported preferences for new runs;
  Advanced exposes individual phase overrides.

The README introduces the current version. This file retains the complete
version history.

**Release-history truth.** This file is prose, and prose drifts: five
releases (v2.5.0, v2.5.1, v2.6.0, v2.8.1, v2.8.2) shipped with no tag at
all and nobody noticed for months, and two rows below name versions no tree
ever declared. The source of truth for declared version trees is the version
the plugin manifests actually held, commit by commit, on `origin/main`; tags
and explicit superseded-candidate dispositions establish which trees were
releases. `scripts/ci_release_tags.py` verifies this table and the tags against
that history and fails CI on drift. Two rows are exempt, each with a recorded
reason:

* **v2.3.0** shipped in the SAME commit as v2.3.1 (the manifest went
  2.2.1 → 2.3.1 in one step). The `v2.3.0` tag points at that shared commit.
* **v2.4.0** never shipped as its own version — its row was added by the
  commit that bumped the manifest to 2.5.0, so its content went out inside
  v2.5.0 and no tag can honestly point anywhere.

**v2.7.2** is absent on purpose: the number was burned during the v2.7.x
lens rewrite and never bumped to.

> **Forward-repair status.** v2.17.20 remains released-incomplete and v2.17.21
> remains the historical source-integration boundary on `main`. The unreleased
> v2.17.22, v2.17.23, v2.17.24, v2.17.25, v2.17.26, and v2.18.0 candidates are superseded;
> v2.18.1 is the tagged local predecessor, and v2.18.2 through v2.18.10 are
> superseded unreleased candidates. v2.19.0 is an unreleased restored baseline,
> and v2.19.1 is a reverted unreleased candidate. The forward candidate is
> v2.25.0. Historical
> graph revision `2757822e` remains an attributed inherited limitation: no
> history rewrite, no re-release of v2.17.20, and no verifier weakening.
> Preparing, validating, or pushing v2.25.0 source is not a tag, upload,
> Marketplace publication, installation, or release claim; those actions
> retain separate human authority.

**Unreleased R-0001 delivery note.** Delivery metrics now have one closed,
redacted, non-cumulative receipt. It binds settings, CI, dashboard publication,
cleanup, portfolio, token/session/worktree, and dispatch evidence by digest;
records the approved suite, feedback, critical-path, parallelism, leak, usage,
Plan-return churn, and end-to-end baselines/targets; keeps billing,
host-observed usage, and archive
upper bounds distinct; names deliberate serialization; and refuses sign-off for
owned leaks or an unexplained hard-ceiling breach. Dashboard, Retro,
Engineering, and release projections consume the sealed receipt without
recounting archived traces, reruns, canceled heads, render output, or DOM state.

| Version | Highlights |
| --- | --- |
| **v2.25.0** | Shared delivery flow, task decomposition, styled dependency graph and dashboard across Claude and Codex; native session and lens usage with advisory telemetry. |
| **v2.24.0** | **Superseded untagged source candidate.** Native review simplification was prepared on main and then restored to the 2.23.1 baseline. The shared advisory flow and cross-host dashboard now continue in v2.25.0. No version tag or archive release was published. |
| **v2.23.10** | **Native review and planning recovery.** Restores authenticated child input and failed-attempt retries, adds whole-repository scope, preserves validation failures and actual reviewer starts, and fixes dependency detection and tool cache recovery. |
| **v2.23.9** | **Shared entry readiness and recovery.** Separates setup from runtime readiness on both hosts, carries Claude startup identity, diagnoses stale engines, and aligns installed-package journey checks. |
| **v2.23.8** | **Native plugin startup and budget recovery.** Uses the host-selected plugin for hooks and review commands, fixes Claude-only packages and failed setup reporting, and resumes existing tasks after an approved token increase. |
| **v2.23.7** | **Native Codex integration.** Native command sessions, permission profiles and preview panels with exact TaskPlane evidence binding. Removes the custom file reader and fixes public command launch defaults. |
| **v2.23.6** | **Shared entry initialization.** Missing setup is repaired once; incompatible read-only file tools refuse activation before a session can be locked. Incident report included. |
| **v2.23.5** | **Session isolation and fresh-checkout readiness.** Each host conversation owns its contracts, run bindings, meters and review state. Native hook proof follows only the same session across checkouts; clearing a review leaves other sessions intact. |
| **v2.23.4** | **Automatic standalone review startup.** Prepares the actual checkout launcher and signed native worker contracts; retries preserve leases and evidence. Fresh installed end-to-end validation remains pending. |
| **v2.23.3** | **Model-led orchestration with strict harness controls.** Disables budget waivers, denies worker control calls, checks fresh hooks on mutations, and safely replays exact phase collection through the existing gate. |
| **v2.23.2** | **Budget and harness enforcement, simpler onboarding.** Applies phase caps during execution, limits each quick lens to 100,000 native tokens and eight actions, meters every tool, and stops budget-triggered continuation loops. Reuses completed reviews, returns findings inline, and keeps onboarding to initial setup. Includes project-local storage recovery and disables harness bypass. Marketplace packages require loading the updated plugin hooks. |
| **v2.23.1** | **Onboarding readiness repair.** Uses the shared hook claim guard for event identity, checks the current launcher after reinstall, and follows the same Git-family path as hook execution. The dashboard keeps incomplete setup visibly incomplete. Every session's first request and every installation/update presents onboarding and preserves the original goal. Prompts hook trust before reload, reports failed acknowledgment writes, and bounds Stop reminders without releasing completion gates. Includes focused regression coverage; a fresh-host check remains separate from CI. |
| **v2.23.0** | **Superseded untagged harness marketplace candidate.** Gives the refactor integrated in [PR #22](https://github.com/vdemkiv/taskPlane/pull/22) a distinct version from the earlier 2.20.0 development builds. Includes stateless phases through Retro, unified lens routing and collection, evolving dependency graphs, dashboard/settings/onboarding wiring, and the cleaned test suite. Package preparation does not claim Marketplace publication. |
| **v2.20.0** | **Superseded local development version.** Product, Design, Plan, Build, Evaluate, EM review and Retro use the v4 run aggregate and verified artifact handoffs. Independent task phases run in isolated worktrees and join at EM review. Dependency graphs can evolve between phases. Singleton execution, cutover switches and unreleased-state migration are removed; prior v3 runs must be archived. Human plan approval and final sign-off remain. |
| **v2.19.1** | **Reverted unreleased candidate.** PR #15 declared this version; [PR #18](https://github.com/vdemkiv/taskPlane/pull/18) restored the exact 2.19.0 baseline without rewriting history. No v2.19.1 tag exists; this candidate is not reused for 2.20.0. |
| **v2.19.0** | **Fresh-task native hook bootstrap and installable package repair — superseded unreleased baseline.** The installed OpenAI package now preserves native SessionStart authority, accepts its launcher-only hook commands during onboarding, and can record the first session-bound receipt in the canonical default store before a fresh linked task has a governed locator. Repository bridges, custom homes, ambiguity, and unrelated chats remain fail closed. An extracted-package regression executes the real onboarding and linked-worktree path end to end. [PR #18](https://github.com/vdemkiv/taskPlane/pull/18) explicitly restored this unreleased baseline; no v2.19.0 tag exists. |
| **v2.18.10** | **Fail-closed delivery authority and isolated global hooks — superseded unreleased candidate.** It completed canonical settings, authenticated native metering, evaluator quality evidence, dashboard-source validation, global-hook isolation, and package provenance, but the installed OpenAI archive still rejected its sanitized launcher-only commands during onboarding and fresh native SessionStart could not record a receipt before locator creation. v2.19.0 closes both bootstrap edges. |
| **v2.18.9** | **Native telemetry, isolated pickups, and one current dashboard — release candidate, not released.** Provider-owned Codex counters are captured at spawn and terminal boundaries, resumed sessions are delta-attributed without double counting, null or zero metering fails closed, and non-zero per-pickup targets and ceilings from canonical settings reach actual hook enforcement. Workers inherit zero conversation turns. Retro and dashboard consumers receive measured totals, while the surfaced dashboard is the full styled current document with a browser-verified visible dependency graph rather than a stale fragment. |
| **v2.18.8** | **Canonical terminal Plan graph fallback — superseded unreleased candidate.** It restored the authoritative terminal task DAG and truthful unverified waves, but native counter enforcement, zero-context spawn binding, and the exact surfaced dashboard document were not wired end to end. v2.18.9 closes those runtime edges. |
| **v2.18.7** | **Terminal dashboard and artifact wiring repair — superseded unreleased candidate.** Terminal `failed` retained the bound Design graph and migrated private artifact classes, but a real run with a drifted mutable Plan file still hid its Plan DAG and waves even though canonical task dependencies remained in loop state. v2.18.8 closes that real-run projection gap. |
| **v2.18.6** | **Truthful legacy terminal migration — superseded unreleased candidate.** It added the narrow fail-closed migration that preserves task status and prior evidence, but the terminal dashboard stage map omitted `failed`, the presentation layer ignored the migrated run's private manifest, and candidate receipts continued to name the baseline instead of the observed delivered revision. v2.18.7 repairs those existing graph, artifact, and candidate-authority edges before upload. |
| **v2.18.5** | **Full R-0002 control-plane wiring — superseded unreleased candidate.** Every new Design run decomposed before dynamic host dispatch; failure classification preceded correction; Build enforced test-quality progression; and durable run-owned artifact classes preserved dashboard, graph, telemetry, activity, validation, cleanup, and Retro evidence. An active run initialized by the previous installed version lacked the new artifact binding, however, so 2.18.5 could not truthfully terminalize it after upgrade. v2.18.6 closes only that migration seam before upload. |
| **v2.18.4** | **Dashboard execution-status repair — superseded unreleased candidate.** The Plan DAG correctly joined immutable Plan structure with governed execution truth, but the broader R-0002 review found that dynamic Design dispatch, decomposition, failure classification, Build-quality admission, and durable artifact retention were not yet wired end to end. v2.18.5 closes those production edges before upload. |
| **v2.18.3** | **Canonical delivery settings, CI-first testing, and exact-owned cleanup — superseded unreleased candidate.** A typed fail-closed loader established the single settings spine and one settings-bound dashboard snapshot, but the final Plan graph retained Plan-time pending task values after the governed loop completed. v2.18.4 corrects that producer/consumer join while preserving the 2.18.3 architecture, evidence-pruned suite, CI policy, cleanup, release provenance, and metrics. |
| **v2.18.2** | **External-worktree Codex hook bootstrap repair — not released.** Codex hooks now resolve the primary checkout's validated launcher from Git's common directory when an external linked worktree lacks its ignored local bridge, route the host-native probe through the same stable engine, fail with an actionable error when no launcher or host plugin root exists, and restore a missing runner before attempting any required hook-config rewrite. Repository-family resolution now uses Git's absolute top-level and common-directory facts so all linked worktrees share the same durable Taskplane state. Exact external-worktree regressions cover PreToolUse, SessionStart, and Stop in both native and bridge manifests. A local reinstall and generated upload artifact are development actions, not Marketplace publication or a public release. |
| **v2.18.1** | **Complete local and exact-PR-head-SHA release proof — tagged local predecessor, not publicly released.** Adds the standard-library closed-inventory local CI runner, deterministic exact-once receipts, isolated parallel shards, deadline-safe cleanup, mutation detection, and a blocking PR-head proof that synthetic merge validation cannot substitute. Its tag and local installation do not claim an upload or Marketplace publication. |
| **v2.18.0** | **Superseded public-metadata correction candidate — not released; compatibility N-1 for v2.18.1.** Removes external orchestration-product comparisons from every public skill and agent discovery description plus public security guidance, while preserving exact external namespaces only in internal collision enforcement and tests. Restores the deleted repository-wide branding regression test so a future merge cannot silently reintroduce the copy. It was never promoted into released truth. |
| **v2.17.26** | **Superseded R-0002 whole-codebase EM remediation candidate — not released; compatibility N-1 for v2.18.0.** Integrated the completed high-, medium-, and low-priority remediation waves, focused final defect fixes, exact-SHA verification evidence, and final Engineering Manager disposition. It was never promoted into released truth. |
| **v2.17.25** | **Superseded R-0013 main-integration candidate — not released; compatibility N-1 for v2.17.26.** Completed the R-0013 Plan: Codex-native execution authority replaced Taskplane-owned delivery lifecycle control; retained Design evidence validated one concurrent all-26 Design-only sweep; Build, Fix, Evaluate, and execution-time EM started zero Taskplane lens workers; deterministic ready sets, one event-driven wait, bounded acceptance waves, fail-closed native usage evidence, real-checkout wiring closure, and atomic exact-SHA terminal truth governed delivery. It was never promoted into released truth. |
| **v2.17.24** | **Superseded local candidate — not released; compatibility N-1 for v2.17.25.** Applied the minimal R-0013 bootstrap correction: W31 producer observations flow through the Codex-native producer-consumer adapter, and execution-time EM uses the governed zero-lens path with an explicit empty collection rather than starting or waiting for a lens producer. It was never promoted into released truth. |
| **v2.17.23** | **Superseded local candidate — not released.** Corrected the hook-home binder so a secure explicit locator may select either a dedicated home or canonical `$HOME/.taskplane`; missing locators, noncanonical locator homes, and inherited `TASKPLANE_HOME` conflicts still fail closed before any receipt. Native hooks, the repository bridge, the generated `.taskplane/codex-hook.py` launcher, and the host runtime retain one locator-bound home without receipt copying. It was compatibility N-1 for v2.17.24, but did not close the W31 producer-consumer and execution-time zero-lens EM bootstrap seams. This row does not claim a tag, upload, installation, publication, push, or release. |
| **v2.17.22** | **Superseded Marketplace candidate — not released.** Packaged the completed two-part R-0001 delivery after the native-dispatch correction, but its bootstrap binder incorrectly rejected an explicit secure locator for canonical `$HOME/.taskplane`. It was never promoted into released truth. |
| **v2.17.21** | **Source-integration boundary.** The complete R-0001 change set was integrated and pushed to `main` at this version: forward release evidence, hosted default-branch preparation, package parity, Design/checkpoint closure, performance telemetry, same-SHA pickup work, and the final removal of duplicate scheduler authority. This historical row does not claim a Marketplace upload or installation. |
| **v2.17.20** | **Stateless pickup now supports explicit, repository-only continuation under attributed human authority.** `tp pickup <design-contract> --trust-source <exact-source-sha>` loads the approved Design Contract and committed shelf evidence from the checkout, records the supplied flag and exact source SHA verbatim in the pickup receipt, and enters the existing direct-assignment/BUILD-C checkpoint path without creating run, track, claim, or lease state. The assertion establishes structural agreement with the selected source SHA; it does not cryptographically authenticate the operator, producer, shelf, or repository origin. Missing or malformed assertions, source mismatches, malformed evidence, structural tampering, and broken receipt lineage fail closed before BUILD-C. The incumbent v1 symmetric shelf handling remains unchanged on its existing private-secret path. |
| **v2.17.19** | **BUILD-C is now part of the governed production flow, not dormant implementation.** AC-bound checkpoints are executed through the live command-event runtime and mint engine-observed, bounded-output receipts tied to the exact environment and revision. Agent submission reaches those checkpoints; only a green checkpoint can authorize integration; DEFINE uses the existing bounded selector; and graph-disjoint direct assignments run concurrently without default claim, build-lease, or wave state while overlaps serialize. The bootstrap path now validates non-empty task-local criteria, preserves only immutable dependency-closed passes across metadata-only replans, binds re-anchor authority to signed structured evidence, loads the review runtime strictly from the target checkout, fingerprints all routed modules including graph quality, refuses symlinks and byte drift, resolves graph quality before routing, activates the selected quick producers before dispatch, and waits once on the event stream. Focused regression coverage pins tamper, stale-revision, nested-false, target-leakage, TOCTOU, ordering, activation, and severed-edge failures. Automatic full, deep, promoted, serial-all, and all-26 review remain prohibited. |
| **v2.17.18** | **Historical R-0007 candidate — not retained on the surviving `main` line.** Commits `5b409b9` through `75bb8d3` implemented and tagged the candidate, but merge `c6800b6` retained its first-parent tree unchanged and therefore did not integrate the R-0007 changes. Do not treat direct human-authorized deep review, the R-0007 production-flow wiring, or its reduced SCC inventory as shipped by v2.17.18. Later versions independently restored or superseded some mechanisms; their own release rows and the current code define availability. |
| **v2.17.17** | **Public Taskplane metadata describes Taskplane alone.** Skill and agent discovery descriptions no longer name or compare Taskplane with another orchestration product, and public security guidance uses Taskplane-native terminology. Exact external namespaces remain only in internal collision detection and its regression fixtures, where they are required for isolation rather than product copy. |
| **v2.17.16** | **Live governed delivery now invokes the mechanisms it advertises.** Governed command launch, reconnect, and event-driven waits reach the durable command runtime instead of remaining isolated helpers. Automatic review selects exactly four or five relevant light-sweep lenses, keeps architecture in the set, runs the selected lanes concurrently, and creates no automatic full, deep, promoted, or 26-lens execution. During host screening, signed review actions activate their exact leased contract before command execution, so current code—not a stale installed engine—can reach the production flow. Focused flow-wiring and mutation coverage pins every new invocation edge. |
| **v2.17.15** | **Compatibility, honest graph boundaries, and quick-only governed delivery are now enforceable release gates.** CI compiles and imports the exact tracked Python surface on CPython 3.10–3.13 before tests, rejects unsupported syntax without mutating governed state, exposes parser degradation in graph output, preserves debug tracebacks while giving normal CLI failures bounded recovery, runs the zero-token corpus without model or network egress, and refuses pushed-green claims unless fetched `origin/main`, checked SHA, and required receipts agree. Quick-only review policy emits no deep slots or promotions and returns substantive findings to the same task. The import-cycle ratchet seals measured SCC membership, edges, and physical size before cuts; the first lower-owner split removes the direct `depgraph`/`decompose` cycle while preserving public signatures, payload meaning, fresh direct-call lens maps, cache and fail-open behavior. Checkout-bound nested Python now consistently imports the governed task checkout instead of stale main modules. |
| **v2.17.14** | **R-0004 makes lifecycle stages independent governed entities joined by bounded, verified handoffs.** Product, Design, Build, Review, Evaluation, Engineering, extension, status, sign-off, and Retro flows now operate on immutable stage lineage and confined execution roots instead of transferring predecessor runtime context. Atomic split transactions give children deterministic identities, dependencies, budgets, artifact subsets, and separate roots; terminal stages cannot reopen, and later reuse requires new authority. Conservative singleton migration fingerprints and retains source records, conserves governed references, and records ambiguity explicitly rather than guessing state. Bounded projections and cross-host rollout preserve text and machine parity while the existing R-0003 evidence, authority, interference, merge, recovery, and exact-worktree fail-closed cleanup boundaries remain unchanged. |
| **v2.17.13** | **Stateless governed review authority now survives fresh workers, exact worktrees, and real installation.** Signed proof-carrying actions let evaluator and lens workers bootstrap only their leased authority without relying on inherited hook state. Managed registrations reconstruct a run-bound workspace locator so graph scans, evidence, and result contracts stay attached to the exact run and task worktree; foreign, stale, dirty, active-contract, evidence-needed, and collision cases continue to fail closed during binding and cleanup. The OpenAI hook manifest now uses only parser-accepted root keys, Codex discovers the fixed bundled host-native contract, Claude retains manifest-based discovery, and both deterministic marketplace bundles include and validate the hook manifest, declaration, and runtime adapter. |
| **v2.17.12** | **ReviewKernel evidence now stays exact, executable, and bounded.** Canonical diff construction accepts an explicit governed path set, retains scoped tracked and untracked task files, keeps an explicitly empty scope empty, and fails visibly instead of returning a successful empty patch when the scoped bound is exceeded. Producer registration filters from durable ready-run summaries before loading state. The v2.17.11 reliability behavior is pinned by focused coverage for blocker-summary normalization, immutable findings and provenance, retry execution lineage, exact approval conservation, evaluator-outage identity, active-run repository status, claimed-worktree binding, unsupported-Python refusal, and platform-correct continuation commands using `python3` on POSIX and `py` on Windows. |
| **v2.17.11** | **Codex host readiness now proves the correct runtime without breaking Claude.** The shared SessionStart declaration remains Claude-native, while Codex workspace installation translates it to exactly one `--host codex` checker on POSIX and Windows, replaces stale host-runtime rows idempotently, and preserves unrelated hooks. Regression coverage exercises repeated installation, rejects stale Claude arguments in generated Codex configuration, and keeps repository-manifest fallback expectations accurate. |
| **v2.17.10** | **Final review and recovery now stay truthful across task worktrees and evaluator outages.** Evaluator evidence binds to the claimed worktree, infrastructure-unavailable results remain non-judged and keep readiness closed, and an explicit retry returns to independent evaluation instead of opening a product fix. Canonical collector normalization repairs only derivable verdict/count metadata, rejects evidence-free pass promotion, and preserves affected-producer retry. Risk-scaled routing artifacts, shape-safe status projections, deterministic stage delivery, terminal sign-off sealing, and UTF-8/progress extraction are reconciled across the governed workflow. |
| **v2.17.9** | **Governed delivery now scales review effort to attributable risk and preserves durable progress across the full local workflow.** One consolidated authorization covers routine execution while mechanical Product, Design, and Plan gates still fail closed; recovery and repository continuity survive sibling worktrees; review routing caps and deduplicates promoted lenses, gives documentation-only and simple low-risk changes exactly one risk-selected deep lens, and retains four mandatory floors for substantive or risky changes. Canonical review evidence supports bounded repair and exact execution binding, convergence and per-lens telemetry are persisted from production paths, and status plus large-artifact delivery use durable host-aware projections across Codex and Claude. |
| **v2.17.8** | **Host-native workflow UX now shares one governed semantic model across Codex and Claude.** Capability-negotiated Picture-in-Picture progress reports observed token usage without inventing estimates; native fan-out and authenticated durable approval surfaces retain stable workflow, task, target, and revision identities; dashboards project accessible deterministic carousels and working detail dialogs; and design, build, and dynamic-review previews run from pinned disposable sandboxes with verified isolation, bounded resources, restart-safe process ownership, and fail-closed teardown. Cross-host packaging consumes one capability declaration and rejects stale, duplicate, and post-terminal projections before rendering. |
| **v2.17.7** | **Governed reviews now preserve structured acceptance criteria, enforce production process-tree isolation, and scope lifecycle contracts to exact worktrees.** Canonical DoR and approval outputs retain structured criteria; dynamic validation blocks direct and descendant explicit-URL pushes or fails closed when isolation cannot be proven; sibling Codex tasks in separate worktrees no longer inherit or replay each other's lifecycle contracts. |
| **v2.17.6** | **Post-fix evaluation is incremental instead of reflexively replaying the whole routed review.** When a prior evaluator failure and its ReviewKernel collection are both sealed, the next Evaluate retry dispatches only lenses that actually failed and records prior passing dispositions as reused evidence. Missing, incomplete, mismatched, or unsealed prior evidence falls back to the complete selective route. The final engineering review remains broad, preserving sign-off coverage while removing repeated no-value lens waves during bounded correction cycles. |
| **v2.17.5** | **Large reviews stay bounded without turning every evaluator retry into lease-management work.** Scoped evidence uses content-bound untrusted-data markers, rechecks the 16 KiB producer boundary, and binds overflow references to the active canonical envelope. Each leased lens dispatch now has a run-unique native task identity, so Codex can create a fresh legal child instead of colliding with a prior task path. Host-observed result writes immediately release only their completed producer contract, and dead pre-publication reservations recover from `ready` without weakening staged publication recovery. Focused ReviewKernel coverage includes delimiter forgery, cross-envelope substitution, exact contract release, unique dispatch identity, and dead-owner recovery. |
| **v2.17.4** | **Review readiness, runtime validation, and the final dashboard now tell one evidence-backed story.** Review opening discovers requirement artifacts structurally: descriptive PR commit bodies become implementation requirements and acceptance criteria, while review-oriented instruction lists are classified separately as explicit lens requests and matched against the live catalog rather than one hardcoded phrase. Requested lenses are promoted to deep dispatch; ambiguous behavioral requirements remain `cannot_verify` until sufficient evidence exists. Canonical collection now dispositions every criterion as `met`, `partial`, `not_met`, or `cannot_verify`, attaching exact diff/file anchors, validation mode, related findings, and gate impact. The dashboard renders separate DoR and requirements-validation panels, preserves historical routing truth for older runs, and no longer invents a human decline when rendering was not selected. Inline widgets carry a self-contained adaptive theme and CSP-safe delegated controls for filtering, expansion, approval, and review actions. Dynamic validation may repair a broken build only inside a disposable push-disabled copy of the PR, records the original failure as a review defect, and never pushes those repairs. Regression coverage pins artifact classification, directive-to-lens routing, criterion verdicts, legacy-envelope compatibility, dynamic consent, sandbox identity, widget controls, and dashboard truthfulness. |
| **v2.17.3** | **Review execution, dashboard delivery, and result repair now form one coherent human-facing flow.** Standalone ReviewKernel preflight recommends dynamic validation, surfaces the exact discovered build/test commands and dependency-install requirement, and requires an explicit human choice before dispatch instead of treating an unavailable local toolchain as permission for a silent static review. Codex widgets use `window.openai.sendFollowUpMessage` while retaining the legacy bridge, and execution buttons send the engine's exact host-receipt prompt. Review payloads expose inline fragments with durable HTML only as fallback, and the already-embedded dependency graph is no longer emitted as a redundant artifact. Collection now validates every leased slot before failing and returns a structured repair batch naming every slot, path, producer, and reason, so all producer-owned summary corrections run together before one retry. Focused regression coverage pins command disclosure, exact approval bridging, inline delivery, single-dashboard output, aggregate repair CLI output, and clean collection. |
| **v2.17.2** | **Completed review publication now recovers safely after a hard process death.** The completed-run collection fast path first verifies that its revision is still canonical, then removes only an exact-run reservation owned by the current process or a provably dead process. Reservations for another run remain untouched, and a live different owner is still rejected. Regression coverage proves the abandoned revision-1 reservation is removed and revision 2 can collect normally. |
| **v2.17.1** | **Remote-source approval and managed lens execution are now engine-owned end to end.** When external run storage is unavailable, taskPlane writes a credential-free bootstrap action in the caller workspace before returning `needs_user`; a repeated start returns the same run/action and cannot continue until an explicit human `review resume` or `repository resume`. ReviewKernel now creates exact leased-result directories before dispatch, resolves native Codex child writes from their host cwd back to the validated managed checkout before containment screening, treats exact `review collect --run-id` as bounded control-plane recovery, and releases only the completed producer slots after every leased artifact has landed. The lens role no longer races collection by self-clearing. Focused preflight, ReviewKernel, Codex compatibility, and a complete managed-PR start/write/collect/signoff journey cover the field failure without re-running the repository-wide suite. |
| **v2.17.0** | **Review collection now treats the sealed leased artifact as the authority and host provenance as independently validated telemetry.** An exact allowed result path, immutable run/slot/lease identities, strict versioned schema, complete routed-lens coverage, and finding/verdict consistency remain mandatory. When a host receipt exists it must still match the dispatched child and exact bytes, but an unavailable receipt no longer throws away an otherwise valid lens result. Every accepted result produces a content-addressed validation record carried into the canonical revision and collection manifest, preserving an auditable distinction between `host-observed` and `leased-artifact` trust. The CI/test boundary is reduced without weakening the authoritative gate: Python 3.12 runs the complete pytest suite; Python 3.10/3.11 and hosted OS legs exercise focused portability seams; one public unittest store-isolation canary replaces redundant partial rediscovery; duplicate/static-policy tests are consolidated and the obsolete unittest-floor inventory is removed. |
| **v2.16.9** | **Managed Codex PR reviews now preserve one canonical identity from checkout through collection.** Review patches start at the pinned merge-base rather than the moving base-branch tip. A complete scanner graph that reaches its explicit local-depth boundary remains visibly policy-limited without being misclassified as degraded, and body-only changes with high-confidence module evidence do not fabricate missing-symbol uncertainty. Hybrid-storage write screening now grants only exact sealed lens-result paths under validated taskPlane run roots, records their host-observed producer receipts, and continues to deny Bash provenance bypasses and unrelated external writes. `review collect` and `review signoff` can resolve an explicit ReviewKernel ID from the parent Codex workspace back to its verified managed checkout. A new workflow eval covers merge-base selection, graph policy, selective routing, external leased writes, provenance, parent-workspace resolution, and single canonical collection while leaving model findings nondeterministic. |
| **v2.16.8** | **Large-repository PR acquisition is now target-bounded, transport-resilient, and resumable without redundant network work.** Managed mirrors are created with `git init --bare` plus an `origin` remote instead of `git clone --mirror`; PR preparation fetches only the requested base branch and pull-request head. GitHub RPC/HTTP 400, HTTP/2 stream, early-EOF, and remote-hangup failures are classified as network transport errors and receive one bounded HTTP/1.1 fallback before an honest retry/cancel prompt—never a misleading credentials prompt. Once a run is ready, its pinned checkout and target evidence are reused without another GitHub fetch; downstream review still verifies the local head and immutable diff. Focused tests cover targeted mirror construction, compatible transport fallback, prompt classification, and ready-run idempotency. The preserved real Backstage PR 35183 failure recovered in place, pinned its exact three-file diff, reopened in 0.38 seconds without network access, and reached a governed ReviewKernel run with 5 deep lenses plus one light sweep. |
| **v2.16.7** | **Codex PR review now uses runtime truth and completes preflight before governance begins.** A bounded, session-aware hook receipt proves that the active native or repository hook actually executed, allowing the native task hook to govern managed PR checkouts without repeatedly asking the user to open another task. Repository acquisition materializes the exact immutable binary diff so partial-clone blob downloads and their authentication/network prompts happen inside preflight rather than surfacing later as false graph sparsity. ReviewKernel identities now include the fallback policy revision, obsolete sparse runs cannot shadow a corrected run, oversized impact data travels by verified envelope reference, and taskplane-generated untracked `.codex/hooks.json` is excluded from both Git status and the canonical review patch while tracked repository configuration remains reviewable. Focused boundary tests and a real Backstage PR 35183 smoke prove three changed files, `ready` status, degraded-graph disclosure, and five selective slots without a task restart. |
| **v2.16.6** | **Repository acquisition, managed checkouts, and generated evidence now share one explicit hybrid storage contract.** Hosted repository identities are stable across HTTPS and SSH forms; source mirrors and worktrees live under `~/.taskplane/checkouts`, project knowledge under `~/.taskplane/projects`, run journals/state/graph/evidence/lens output/artifacts under `~/.taskplane/runs`, and reusable graph results under a repository/head/scanner cache. A Git-metadata locator binds each managed checkout to those roots without placing source inside review artifacts. Repository and PR preflight creates an inactive resumable run before acquisition, verifies the requested target, and turns missing tools, authentication, network, storage, or uninitialized local repositories into bounded in-chat actions that resume the same run. Managed parallel tasks receive isolated external worktrees and child run roots. Standalone Review no longer confuses incomplete graph enrichment with an unusable immutable PR diff: it routes selectively from the sealed diff with a visible degraded-graph warning and architecture/security floors, while Evaluate and final Engineering preserve the strict `impact_incomplete` zero-dispatch gate. Runtime projections, dashboards, graph visuals, evaluation contracts, ReviewKernel state, and submission evidence resolve through the same storage locator; legacy unmanaged workspaces retain their existing paths. Scenario fingerprints and storage/preflight/review regressions pin the new flow. |
| **v2.16.5** | **Pull-request review now fails before expensive or stale work when the checkout cannot prove the requested target.** A bounded preflight verifies repository identity, base resolution, merge-base availability, checked-out head, and a non-empty diff before graph scanning, contract activation, or ReviewKernel state. Shallow clones receive one bounded `git fetch --deepen=256`; unresolved history returns a structured `merge_base_missing` refusal instead of silently reviewing an incomplete range. Target fingerprints bind origin, base, head, merge-base, and shallow state, while review cache identity also binds the graph revision, preventing a moved PR head or refreshed graph from reusing stale routing and blast-radius evidence. Standalone review distinguishes target failures from sparse graph evidence; Evaluate and final engineering retain their existing `impact_incomplete` contract. The Windows hook guard now classifies conditional launcher commands by their invoked taskplane action, preserving fail-closed enforcement without mislabelling lifecycle hooks. Focused target, refusal, routing, and guardrail tests cover the release. |
| **v2.16.4** | **Codex hook execution is no longer coupled to a disposable version cache path.** Bundled native hooks prefer the repository's stable `.taskplane/codex-hook.py` launcher and retain the loaded native hook identity for exactly-once processing. The generated launcher stores only its marketplace installation family, validates contained candidates against the taskplane manifest and strict numeric semantic version, rejects malformed identities and symlink escapes, and resolves the newest valid engine on every call. All primary skills prefer that same launcher, with the loaded plugin root retained only for first setup and other hosts. The first upgrade from a pre-resolver hook still needs one task boundary because already-cached command bytes cannot be changed retroactively; subsequent engine updates do not require a Codex restart. Changed skill or MCP definitions remain host task-bound. Focused compatibility, Windows-manifest, onboarding, documentation, release-freshness, and security regressions cover the release. |
| **v2.16.3** | **Release and Codex onboarding regressions are closed.** The README release window is rotated inside the release batch so its exact-three-row CI contract cannot be broken by the final version commit. Missing or unsafe language-reference files and named sections now produce a structured, zero-dispatch `mapper_unavailable` result on both selective and legacy routes instead of an uncaught traceback, preserving the documented fail-closed boundary. Codex repository-hook onboarding retains `TASKPLANE_HOOK_PATH=bridge` in every generated POSIX and Windows command, keeping native and bridge receipts distinct for exactly-once lifecycle processing. Focused regressions cover all three seams, and the release uses one CI push. |
| **v2.16.2** | **Host-capability CI and workflow contracts are synchronized without changing the shipped feature surface.** The per-task cost ratchet now requires exactly three gates—two gates is incomplete execution, not an optimization—while the canonical loop fixture once again exercises that full path. The 16 host-capability receipt variables are documented; ReviewKernel workflow tests consume leased slots and strict evaluator output rather than legacy `deep`/`sweep` or free-form findings shapes; runtime-evaluation fixtures provide the required final engineering facts; and canonical cross-host comparisons normalize host-local route receipts without deleting them from dispatch evidence. The pytest-only manifest includes the three new capability/evaluation suites. `loop.py` remains below its existing 3,100-line ceiling by extracting status projection into `loop_status.py` instead of raising the pin. Local verification: 2,867 passed, 2 skipped, 865 subtests; the Linux/Python 3.10–3.12, macOS, Windows, docs, packaging, Codex hooks, dispatch parity, and release-history GitHub Actions matrix is green. |
| **v2.16.1** | **Model-evaluation availability is guidance, not a fabricated product failure.** EVALUATE gains a structured `unavailable` outcome for bounded host, transport, agent-lifetime, and producer-receipt failures. The gate admits it only when the execute/fix suite is green, no acceptance criterion is not-met, no completed lens blocks, and the evaluator records a versioned reason. It advances with a visible warning, preserves human sign-off, never increments product fix cycles, and refuses an incorrectly submitted `fail` for the same pure outage. The evaluator brief, strict schema, CLI, frozen dispatch contract, loop status, and canonical harness all carry the same rule. The accompanying two-session retrospective records a value-delta prerequisite, proportional validation matrix, bounded model-eval budget, canonical-run reuse, and self-hosting circuit breaker for follow-up enforcement. |
| **v2.16.0** | **Crash-safe Retro and a leaner, more portable validation boundary.** Retro is now a dedicated resumable module: it reserves a stable run identity before side effects, deduplicates knowledge and trace records across retries, verifies its receipt before closing the loop, and remains recoverable after interruption. CI fixtures now exercise the current Product DoR and human-gate contracts without weakening production gates; subprocess encoding and Windows checks are explicit; the unittest inventory is honest; and the cost ratchet remains at one suite execution. Runtime cleanup closes dashboard/knowledge file handles, removes redundant review blast rendering and stale private helpers, and unifies duplicated evaluation status vocabulary. This is the same approved v2.15 workflow surface with reliability, portability, and maintainability fixes. |
| **v2.15.0** | **Approved skill-flow contracts and convergent review.** All ten user-facing skills now ship a portable, human-approved `flow.json` that pins stage order, graph use, worker action, visualization, and the gates that genuinely require a person. The mission-control dashboard renders workflow execution and dependency impact in the established taskplane report hierarchy; intermediate transitions defer delivery/acknowledgement until a human gate instead of repeatedly asking for approval. Product now has one mechanical DoR shared by standalone Product and the PM stage, requires security/architecture statements, links requirement scope into the graph, and records explicit approve/change sign-off before Build. Review now applies a structural admissibility bar at canonical collection: a defect needs trigger/outcome/repro, a non-functional violation must resolve to an adopted requirement, decision, config, budget, or language-reference section, and everything else becomes a durable non-gating note. Settled identities—including `not-a-defect`—travel to lens briefs by bounded artifact reference and cannot recur without named new evidence; a frozen two-pass scenario pins zero new admissible findings after settlement. The model-evaluation harness fingerprints the approved graphs, observes every nested Codex command, and grades host-specific stop points rather than the wrong workflow. Targeted CI repairs cover portable context paths, setup-phase write exclusion, language-reference acknowledgements in real leased results, regenerated dispatch/stage fixtures, workflow-free gate traversal, and the honest three-gate cost harness. |
| **v2.14.2** | **Language-specific references now travel through routing and canonical review context.** Routing resolves Go, Python, and TypeScript references from the canonical changed-file set, carries the applicable reference by path in both legacy and selective Review briefs, and exposes language-correct priming in serial Execute/Fix action payloads instead of assuming Python. Go Design receives a dedicated package-boundary and concurrency reference, Go 1.26 is the adopted baseline while 1.27 remains draft, and the Python/TypeScript gates correct their previously inert PCRE and typed-lint checks. Windows CI is now a required, focused host-boundary check rather than a second 27-minute full-suite copy. |
| **v2.14.1** | **Codex marketplace installs now establish the enforcement boundary they depend on.** Codex plugins supply taskplane's skills, while onboarding installs a portable repo-local `.codex/hooks.json` plus an ignored machine-local bridge to the installed engine; readiness stays false until that bridge is current, and plugin upgrades are detected as stale instead of silently using an old hook runtime. Review routing now reads one bounded canonical diff with nearby hunk context—not unrelated whole-file content—so security retains enclosing authorization signals while selective routing avoids false-positive lens fan-out. Routing input is capped at 200 files and 64 KiB per file, and large shared review envelopes deduplicate repeated file/symbol/impact data before scoped views are sealed. |
| **v2.14.0** | **Governed review is now one dependency-aware, selective evidence pipeline across Review, Evaluate, and final engineering sign-off.** Graph quality and a bounded blast radius are established before routing; all 26 lenses receive an explicit `deep`, `light`, or `n/a` disposition, architecture and security floors are preserved, and only the deep set plus at most one light sweep executes. Insufficient graph evidence stops at `impact_incomplete` with zero dispatch. A single immutable envelope carries the diff, impact, graph quality, requirements, contracts, DoR, DoD, and runnability; reviewers consume deterministic scoped views by reference and write fingerprint-bound leased results into one canonical findings revision. Evaluate cites matching execute-gate evidence instead of rerunning it, while structural/token counters and the frozen PR-9464 oracle protect both efficiency and the known blocking regression. The dashboard now renders the dependency graph, workflow progression, and approval/rejection gates in taskplane's report visual language. Claude and Codex share byte-equivalent canonical artifacts; host lifecycle hooks are the provenance boundary and must be loaded by starting a new task after updating the plugin. |
| **v2.13.0** | **The budget counted tool calls; the cost was tokens — and the render obligation was the single most expensive thing in the product.** One measured review: 3.77M effective tokens over 99 turns. Four mechanisms, each aimed at a line in that breakdown. **450k went to inline dashboards**, because v2.9.0 made showing an artifact enforceable and the only compliance path was pasting the engine's HTML back through a widget tool — ~52k characters re-authored at output weight, of a file taskplane had already written to disk. `tp findings` now hands back a PATH above `TASKPLANE_INLINE_MAX` (24k chars) and `tp ack <id> --delivered <path>` discharges the obligation with the SAME bytes and the same fingerprint check; a delivered substitute is still recorded as a mismatch. **777k went to 69 shell commands**, about ten of them before a single lens saw the diff: `tp review start` establishes tools, target pin, graph, impact, contract, obligations, routing, runnability and the ready-to-dispatch briefs in ONE payload, deciding nothing — the same move `tp loop evidence` made for the evaluate step in v2.6. **754k went to four lens agents, "each carrying its own copy of the diff and the blast-radius brief"** — the diff is identical for all of them, so it is written once to `.em-review/context/` and the briefs cite the path with an explicit do-not-re-derive instruction. **And the action ceiling could never be tuned:** an action cost ~11,261 effective tokens on that review with a two-order-of-magnitude spread, so raising 40 to 80 bought ~440k tokens sight unseen. Contracts now take `--max-tokens`, read from the host's own transcript and weighted the way cost falls (cache reads ×0.1, cache writes ×2, output ×5 — the same review was ~22M raw and ~3.8M effective). It fails open in every direction — no transcript, a torn line, a missing usage block — because a budget that blocks when its instrument breaks makes a broken instrument into a broken product, and the action ceiling still stands underneath. 2,048 tests, 10 mutations observed failing. |
| **v2.12.0** | **A review can now prove what it reviewed — and `git`/`gh` are dependencies, not conveniences.** Two field reviews of `aws/karpenter-provider-aws#9464` both cloned the repository and neither could prove it: `tp new` took the target as free text and the contract recorded no origin, no base, no head, no record of how the code arrived. Both reports stated the workspace and the diff base in PROSE, by hand, and a review conducted entirely from a rendered web diff would have produced identical artifacts and an identical gate. `tp target` closes it: `fetch` acquires a pull request with the same two git commands every time and records them, `pin` reads what the checkout actually is (origin, head, base, merge-base, dirty paths) and reduces it to a fingerprint, and `tp new --target <pr> --fetch --base <ref>` does both at activation and writes the pin into the contract. Findings cite that fingerprint in `meta.target`; `tp findings` says UNBOUND when they cite nothing or a different tree, and the PreToolUse screener refuses `dod`, `loop submit`, `loop approve` and `loop retro` on a read-only contract until the workspace is pinned — the same obligation-to-prohibition conversion as v2.9.0, never blocking the review itself. **And `gh` is now a declared dependency.** A clone carries the code and none of the intent: a PR's title, body, linked issues and review conversation are not in the git objects, so in the field that context arrived over unauthenticated web reads nothing recorded. `tp onboard` and `tp target tools` report git and gh with versions and auth state, `tp target tools --install` installs gh through the host's own package manager, and a remote-PR review without gh fails loudly instead of quietly degrading. taskplane deliberately does NOT download and execute a release tarball — a hardcoded checksum nobody maintains is a worse guarantee than the package source the user already trusts, and a test pins that target.py never reaches for curl, urlopen or tarfile. 2,013 tests, 9 mutations observed failing. |
| **v2.11.0** | **The applicability engine was never wired to the wave — and three v2.10.0 claims were not true.** Route v2 (content, graph and requirement signals producing a per-lens `deep | light | n/a` verdict, every skip carrying machine-checkable negative evidence) shipped in v2.4.0 and was unreachable from the CLI for six releases: `route()` enables it only when `stage` is passed, `cmd_lens` passed nothing, and the one caller that did pass `stage="review"` was the coverage REPORTER. The engine scored the diff for a report and the wave ignored it. Compounding it, `tp-engineering/SKILL.md` mandated `--all` on every review command, and `--all` disables the engine by construction — two independent causes, so fixing either alone changed nothing. On a Go type change plus a docs edit the glob router summoned 6 lenses deep and marked none n/a; the engine routes 2 deep, 4 light, 20 n/a. **And v2.10.0 claimed three fixes that had not landed.** `graph impact` still could not see intra-repo Go: the root-module prefix stripping went into the JavaScript resolver, the Go branch was never touched, and the helper could not have worked anyway — it looked the root path up by key in a SET, which is what the scanner holds. The three "minor" papercuts were reported done and all three reproduce. **Then the exit path, from a second field run.** `session-verify` demanded `tp ack <id>` while the budget refused `tp ack <id>` — twelve firings, no reachable state that satisfied it; `ack` is now unmetered (it discharges an obligation and cannot widen scope — unlike `clear`, which stays walled), the last actions of a budget are RESERVED for closing rather than added to it, and the hook names the real blocker instead of repeating an impossible instruction. `tp ack <id>` in the wrong directory returned `acknowledged` for an id nobody issued, and `graph html` there emitted 5,684 bytes of valid-looking dependency graph for a workspace that had never been scanned — both now refuse and say where the contract actually lives. `tp init` writes `.git/info/exclude`, never a reviewed repo's `.gitignore`. A read-only contract can create the directory it authorizes. `python3 -c` is screened as a grammar instead of refused as a blob — an allowlist of stdlib reads passes, every write shape and anything unparseable still fails closed. 1,974 tests, 14 mutations observed failing. |
| **v2.10.0** | **Nine defects a real upstream repo found that no self-review could.** v2.9.0 was run end-to-end against `aws/karpenter-provider-aws#9464` in a separate session. The harness held — not one write reached reviewed source — and nine defects surfaced that only appear at somebody else's scale. **The worst was silent:** `tp findings` printed `0 high · 3 med · 13 low` while the findings file carried `class: regression`, which the engine's own gate blocks. The reviewer read the headline, reported "0 confirmed regressions", and recommended approve. The engine had the answer and the headline did not say it — the skill has mandated that split since v2.3.1. The headline now reads it off `loop.classify_findings` (never a second implementation) and says `1 BLOCK (1R·0H·1P·0O)`. **Parallel lens dispatch did not work with more than one agent:** six lens contracts each allowing `.em-review/lens-<id>/**` intersected to the EMPTY set, so 4 of 6 lenses produced no evidence at all and the wave board read 2/6 for a finished review. Sibling waves — every member read-only, every write-allow under one common root — now merge their write-allows and sum their budgets instead of taking the minimum, which had given six agents one agent's 30 actions and killed them mid-task. Contracts that genuinely compete over separate trees still intersect to nothing. **The graph could not see intra-repo Go at all:** a root `go.mod` was skipped as "describes the repo, not a module", so every `pkg/**` import landed as `ext:` and `graph impact` reported 2 modules and no call structure on a 256-module repo — the review's only blocking finding had to be traced by hand. The root module path is now consumed as a PREFIX (never a module id, which would collapse the repo into one node). **`tp version` was broken on every Claude-side install**, including the shipped `.plugin`: it read only `.codex-plugin/plugin.json`, which the Claude package does not contain. CI never saw it because CI inspects the source tree, where both manifests exist. Also: `tp contracts` names stale slots a union is silently applying, `tp clear --all/--slot` releases them from outside, `tp findings --html` emits a self-contained document for hosts that cannot render fragments inline, and `> /dev/null` is no longer screened as a write. **Six lens agents each burned actions rediscovering that `go test` could not run** — one fact about the CHECKOUT, paid for six times, and 41.5% of that session's tokens went to sub-agents. `lens dispatch` now probes build/test runnability ONCE, before composing briefs, using a bounded cheap subcommand (`go list ./...`, `cargo metadata --offline`, `import pytest`) and never the suite itself; the verdict is stated in every brief with an explicit do-not-re-probe instruction, shown on the wave board, and returned as `runnability.summary` for the findings `meta.tests`. It is cached per tree state and PATH, so a whole wave shares one answer while installing the toolchain mid-review re-probes. It is information, never a gate: no screener, contract or gate may read it, and a test pins that. `TASKPLANE_RUNNABILITY=off` skips it. And an interrupted wave no longer costs the whole fan-out again: `lens dispatch --resume` re-briefs ONLY the lanes with no findings.json, reading the same on-disk source the wave board already reads (a corrupt or listless findings file counts as NOT landed, so it re-runs). In the session that produced this report, four of ten sub-agent transcripts existed because a wave was spawned in a turn that died before the agents reported — about 16% of its effective tokens, spent twice. Opt-in by design: a fresh review of a changed diff must re-run every lens. **One request was refused:** the report asked for `tp clear` to be exempt from metering so an exhausted agent could release itself. Clearing leaves the workspace ungoverned, where the screener abstains — that trades a deadlock for a bypass, and `TestTheWallHolds` pins it. Only pure reads are exempt; recovery happens from outside. 1,905 tests; 13 mutations observed failing, one of which found a position bug in this wave's own code (`echo tp clear` read as a release command). |
| **v2.9.0** | **An obligation is now a prohibition — the first mechanism in this product that makes a required artifact non-skippable.** A hook can DENY an action; it cannot COMPEL one. That asymmetry is why every prohibition here has always held at 100% — the screener refuses an out-of-scope write, `rm -rf .`, an interpreter escape — while every OBLIGATION the flow defines ("render the wave board", "show the graph") held at 0%, because the only thing standing behind it was an instruction in a skill. Five structural attempts to close that shipped between v1.5.3 and v2.8.2 and the same complaint arrived after each one. An instruction is not a mechanism. **The conversion is:** not "you must show the graph" but "you may not declare the work finished until the graph has been shown". A conclusion is a command, a command is a tool call, and a tool call can be refused. `tp new --owes review` records the artifacts a run owes BEFORE any of the work starts — so a skip is a recorded fact from the first second rather than an absence nobody can see — and the PreToolUse screener then refuses `dod`, `loop submit`, `loop approve`, `loop retro` and `loop gate` until each is discharged. It is deliberately narrow: it can never block an edit, a test, a search, or any other part of DOING the work, only the act of declaring it done, and `TASKPLANE_OBLIGATIONS=off` disables it while still recording, because a governance mechanism with no stated way out is one people route around by uninstalling. **A second premise fell with it.** `obligations.py` stated that the engine "CANNOT see whether a rendered artifact was actually put in front of a human" because the render "happens in the host, outside every process taskplane runs", and therefore that showing something could only ever be a CLAIM. That was wrong: a PreToolUse matcher is a regex over TOOL NAMES, and MCP tools are named `mcp__<server>__<tool>`, so a `mcp__visualize__.*` matcher reaches the render at the same seam that already screens writes and dispatches. `tp screen-render` now records every render as a FACT with its content fingerprint, which separates three failures that were previously indistinguishable: SKIPPED (demanded, never rendered), SUBSTITUTED (rendered — but not the artifact the engine built, which makes the byte-for-byte render contract checkable for the first time) and CLAIMED ONLY (acknowledged with no observation behind it). Rendering the engine's exact bytes discharges an obligation on its own, so the honest path is also the short one. The engine still cannot read the ledger — the block lives in the screener and the hooks, the deletability contract is unchanged, and `loop.gate` is untouched. A `Stop` hook (`tp session-verify`) reports anything still owed as the net underneath. 1,839 tests; 17 mutations of this behaviour observed failing. |
| **v2.8.2** | **The review's own surfaces: readable, inline, and honest about what they omit.** Found by running the full engineering review against `backstage/backstage` — 12,042 files, 265 packages, seven governed lens-agents in one parallel wave — and then reading the result the way a human does. The analysis was sound; three of the four complaints were about the surfaces. **The clean list was a wall of prose:** `tp findings` joined 35 clean checks with `"; "` into one unbroken paragraph, each sentence itself full of semicolons, and showed only the first twelve under a header that said thirty-five. It is now one row per check with the domain lifted into a label, and the omission names itself and says where the rest lives — a dashboard that quietly truncates its own coverage reads as "that's all there is". **The graph could not be shown inline, only linked:** `tp graph html` embeds every module and every edge, which on a monorepo is a 620 KB page — a fine file and an impossible widget, so the graph kept getting narrated instead of rendered, which is precisely the substitution the obligation ledger exists to catch. `--focus N` crops the map to the changed set plus everything within N dependency hops (620 KB → 30 KB), and `--fragment` carries that page **byte for byte** into an embeddable sandboxed iframe, gzip+base64 so it fits a widget (30 KB → 7 KB). Byte-identity is the whole point: a wrapper that re-authored the page to fit inline would be the same substitution wearing the engine's name — so the fragment is tested by unpacking it and comparing against the original document. Recorded and not fixed, in the v3 Phase 3 backlog (WS-G, since retired): a review still reports no token or cost accounting, an obligation catches a skipped RENDER but never a skipped COMMAND, and the scanner minted 714 modules from a 265-package repo — including one called `${{ values.name }}`. |
| **v2.8.1** | **An evals layer: does the plugin actually get USED, not just work?** Every test in this repo asks whether the machinery is CORRECT. None asked whether it was used — and that is the gap the product kept falling into: "no inline dashboard visualisation, no report nothing", "this is not the graph and dependency visualisation we designed". In each case the engine rendered the artifact, wrote it to disk, pointed at it in the payload and told the assistant to show it. Every unit test passed, because nothing was wrong with any unit. The engine could not see the failure, so nothing recorded it. It records it now. The engine writes down what it demanded (`taskplane/obligations.py`), `tp ack` discharges it, and `scripts/ci_evals.py` scores six areas — artifact surfacing, the product's own graph, agent fan-out, skill-flow order, gate discipline, cross-host parity. The rule that keeps it honest: **only what the engine cannot observe needs a claim.** Fan-out already had expected-vs-observed dispatch behind the PreToolUse Task hook, and step order and approval attribution are already trace events; an obligation for any of those would have been a second record of the same fact, free to disagree with the first. So exactly two kinds need acknowledging — the dashboard and the graph — because `show_widget` happens in the host, outside every process taskplane runs. Two properties carry the design. THE ABSENCE IS THE MEASUREMENT: an obligation issued and never acknowledged is the complaint above, in a ledger, countable — and issuance blocks nothing, because the moment it could cost a gate people would route around it. A SUBSTITUTE IS NOT A SKIP: each obligation carries the engine artifact's content fingerprint, so an assistant that draws its own chart has nothing to cite, and that failure is counted apart from skipping. Measured on a real session, artifact surfacing went **0% → 100%** once the driver skill acknowledges what it showed; the 0% is the honest baseline the layer exists to have made visible. An `evals/` corpus of four profiles — including both complaints verbatim as data, plus an honest-unknown profile pinning that a session is never scored 0% for something the engine could not see — proves the scorer without a host. Instrument contract, same as the yield meter: gates nothing, cannot fail CI or change a verdict, engine never reads the ledger (pinned), deletable with no effect. Also fixed: a test of mine that read ambient graph state and so passed under pytest while failing the `unittest discover` leg, where all 1,500 tests share one process — a test whose result depends on what ran before it proves nothing about the thing it names. 1,762 tests, verified on pytest, on unittest discovery, and on the Python 3.10 matrix floor. |
| **v2.8.0** | **The dependency graph now describes the codebase it reviews — measured, not asserted.** taskplane had a cost meter and a yield meter and no ACCURACY meter for the artifact every review depends on, so the graph drifted a long way without anyone noticing. `scripts/ci_graph_accuracy.py` scores the real scanner against hand-authored ground truth over four repo profiles — polyglot, npm+go workspaces, a markdown plugin, a Java monolith — in the graph's OWN model (node kinds, and the two edge families a caller must not confuse), with each expected edge carrying a DIFFICULTY so the headline is three numbers rather than one: what a scanner should already get, what a manifest reader adds for free, and what must be recorded rather than scanned. It gated nothing and immediately found four defects. **(D-0007) A repo names its own modules; nobody read them.** The workspace profile scored 0% module recall AND 0% precision while missing nothing — four modules found correctly under four names nobody uses (`ui`, `core`, `svc/gateway`) where every manifest, import statement and human says `@acme/ui`, `@acme/core`, `acme/gateway`. An id exists to be cross-referenced; "impact touches ui" is not an answer a reviewer can carry back to the codebase. The rule adopted is narrower than "read every manifest": take a declared name only when that name is what other code IMPORTS the module by — package.json `name` and go.mod `module` qualify, pyproject's DISTRIBUTION name and pom's build coordinate do not, and a ROOT manifest describes the repository rather than a module in it. The refusals carry the design: a rule that swallowed every manifest would rename `services/pricing` to `pricing` and the whole repo to `shopfront`. The map is PUBLISHED in `meta.module_ids` because half this fix is worse than none — impact, completion, decompose, lens routing and the hub signal all resolve changed files the same way the scan did, or they look up `packages/ui` in a graph that only knows `@acme/ui` and report an empty blast radius while looking healthy. **(D-0018) A nested source root corrupted identity two ways at once.** The rule read only the segments AFTER the last source root, which is right when that root is the repo root and wrong in the most common JS and Go layout: `web/src/App.tsx` became `src` — an id naming a CONVENTION that every repo produces — and `web/src/cart/X.tsx` became `cart`, MERGING it with `admin/src/cart/Y.tsx`, so an edge landing on either was reported against both. A generic source root is now invisible to identity, dropped wherever it appears instead of shifting the answer, which makes whether a project uses `src/` stop changing its module ids. The entire 1,679-test suite passed with this defect in place and passed unchanged with it fixed — not one test put a `src/` inside a project directory. The corpus is what found it, and that is the argument for the corpus. **(D-0016) The product is not only source code.** `CODE_EXT` decided what EXISTED, not just what got parsed for imports, so a repository whose product is markdown and declarative JSON was invisible: this plugin's own skills, agents and lenses were absent from the graph AND from decomposition, which is what a review walks to know where to look. Artifacts (.md .json .yml .yaml .sql .tf) are now modules WITH their files. Two boundaries are drawn on purpose — BUILD DESCRIPTORS stay out (a pom says how code is assembled, it is not a thing the code depends on; admitting `.xml` minted a `main/resources` module out of a Java DI file) and so does a root-level artifact, which would land in the catch-all `(root)`. **(D-0015) The graph modelled imports and nothing else.** On a codebase whose components talk by NAMING each other — a skill dispatching an agent, an agent applying a lens, a module reading a routing catalog — "what depends on this?" had no answer for over half the tree: 4 internal edges for this repo. References now resolve to dependency edges, and the load-bearing property is that this is RESOLUTION rather than pattern-matching — the regex is a sieve that decides nothing, and a token becomes an edge only when it resolves to a file that is really in the tree, so prose about a file that does not exist produces nothing. A code file's reference counts only when the target is an artifact; code-to-code stays the import scanners' job. The kind (`calls` vs `uses`) is the one place a convention decides anything and is confined to a LABEL: both live in the dependency family, so grading one wrong changes how a relationship reads and never what a blast radius contains. Measured, corpus-wide: easy edges 33% → 100%, manifest 40% → 100%, declared 0% → 100%; all four profiles now score 100% module recall AND precision. This repo went from 6 modules / 4 internal edges to 28 / 120. Ground truth was revised twice, on the record and with the reasoning stored in the fixtures — once raising the bar (5→6 expected modules) and once reclassifying an edge a root pom cannot attribute to a package into a new `not-derivable` band. **Review integrity, from the same whole-codebase review.** The deep-dispatch budget was absent from `--all`, which is what a whole-codebase review runs: 26 subagents under a cap of 8, so the most expensive review the product performs was the only one with no spending limit — the cap now DEMOTES rather than drops, and an architecture full pass is exempt. Content detectors fired on this repo's own documentation — 17 lenses on a five-file docs change, `dba` DEEP because a doc explains query patterns and `data-safety` on the privacy LENS DEFINITION, whose job is to describe migration markers — so prose is now scanned only for a lens whose own globs claim it (17 → 2, and the review cost mean fell 7.55 → 7.05 with no coverage lost). The suite cache carried a timestamp nothing read, so a green result from months ago discharged today's `tests_pass`; citations are bounded in time and DISCLOSED where the human signs off rather than only in the trace. Two of three cost pins were hand-tallied constants that could not have detected a cost increase while being cited as evidence cost was flat; both now read the engine, and report 11 and 3 where the old numbers said 10 and 4. Release archives name their source commit and refuse a dirty tree, because determinism proved an archive reproducible FROM SOME TREE and never said which — it deliberately does NOT claim CI passed, since a local build cannot know and implying it would launder an unverified artifact. Rendering the dashboard lived inside every state transition under `except Exception: pass`, which is not fail-open but fail-SILENT: the payload TELLS the human the view was "refreshed for this transition", and when rendering threw the key was simply absent — no error, no trace, no warning, and nothing left behind to diagnose. It moves out to `views.py` (the extraction loop.py had been recording as debt since v2.3.0; the dashboard→loop cycle is now genuinely broken rather than smuggled into function bodies, loop.py 3,005 → 2,907 lines) and failures reach the payload, the trace and stderr. Review artifacts were auto-committed into the in-repo team store on every gate transition while PRIVACY.md said they stay local on both plans and that publishing is "a deliberate human act" — model-authored prose is now withheld unless `TASKPLANE_PUBLISH_REVIEW=1`, and what was withheld is NAMED. Trace rotation moved to a fixed `.1`, so the second rotation destroyed the first archive while writing "earlier events moved aside, not lost"; archives are now a monotonic sequence claimed with O_CREAT|O_EXCL. Also: four verified screener bypasses closed (one paren defeated every screen, including read-only), an unscoped contract could overwrite its OWN contract file, `rm -rf .` passed a scoped contract, and four argv-parsing gaps; four regressions in the yield meter itself; and a repo can now declare which trees are not its product code, which took this repo from 28 phantom-inflated modules to 6 real ones. 1,734 tests, and every fix in this release was mutation-tested — the mutation observed failing before the assertion was kept. |
| **v2.7.4** | **Windows portability — all four classes closed, and each one made reproducible on Linux.** taskplane had never actually worked on Windows, and CI could not say so: the failures lived only on the advisory `windows-latest` leg, which is slow, `continue-on-error`, and times out mid-suite — the last run reported 90 failures out of only 645 tests it managed to execute, and blocked nothing. Each class is now fixed AND pinned by a test that reproduces it on the runner you already have. **(1) Encoding.** Every `open()` without an explicit encoding and every printed arrow used the host's default codec — UTF-8 on Linux, cp1252 on Windows — so `tp kb migrate` and `tp northstar` died mid-print, and a requirement file containing `→` could not be written at all. A C locale gives Python an ASCII default, strictly narrower than cp1252, so the whole class reproduces on ubuntu in two minutes and now GATES: 314 failures and 23 errors before, 0 after. 587 call sites gained an explicit `encoding="utf-8"`; the CLI and every script reconfigure stdout/stderr; `ci_unittest_floor.py`'s embedded discovery program is ASCII-only because argv is encoded with the FILESYSTEM encoding, so one em dash in a comment made the spawn itself fail. **(2) Separators — a governance defect, not a cosmetic one.** `norm()` compared a backslash path against a `/`-terminated base with a string prefix test, so EVERY path in a Windows workspace resolved to `ESCAPES:` and the contract screener refused a worker's own in-scope file. Governance that fails closed across an entire operating system is still governance that does not work. Separators are normalized once at each boundary, repo-path arithmetic moved from `os.path` to `posixpath` (`os.walk` yields host separators, `git ls-files` yields `/` — the two enumerations now agree, which is why the git path always worked and the walk did not), and a static ratchet pins the count of host-shaped path calls in `depgraph.py` and `decompose.py` so the class cannot creep back. **(3) Newlines.** Detector reads normalize CRLF, so a file checked out with Windows line endings scores identically to the same file with LF — otherwise one diff routes differently on Windows than in CI, and the byte-frozen goldens are what disagree. **(4) Read-only teardown.** git marks `.git/objects` read-only and Windows refuses to unlink a read-only file (POSIX only needs a writable parent, which is why this is invisible here); `rmtree` clears the bit and retries once, patched in one place rather than at 48 call sites. The new `test_windows_portability.py` drives Windows-shaped input through the code that used to mishandle it — including the complement case, that containment still REFUSES a real escape on a Windows-shaped host. 1,549 tests under UTF-8, 1,547 under a C locale. **Second pass, from the CI leg's own evidence — 90 failures became 6, and the last six named a worse bug than any of them.** `os.kill(pid, 0)` is not a liveness probe on Windows: CPython maps signal 0 to `CTRL_C_EVENT` and calls `GenerateConsoleCtrlEvent`, so taskplane's orphan-contract check SENT Ctrl+C to the console process group. That is what had been truncating the Windows leg — every run died with `KeyboardInterrupt` partway through and reported a partial count that read as slowness, which is why the leg never once finished the suite. It also never measured anything: a dead pid raises a generic `OSError` there rather than `ProcessLookupError`, which the handler read as "unknowable, assume alive", so an orphaned contract was NEVER auto-released on Windows and a workspace stayed governed by a process that no longer existed. Liveness now uses `OpenProcess` + `GetExitCodeProcess`, failing toward governed on anything ambiguous (access-denied means the process exists and belongs to another user, which is not an orphan), with the Win32 branch unit-tested off Windows via an injected kernel32 — including a test that no signal is sent at all. The rest: `.gitattributes` pins LF so byte-identity goldens stop being CRLF-corrupted by the checkout rather than by any code; child processes get `PYTHONIOENCODING=utf-8` and 70 subprocess decode sites get `errors="replace"`, closing a failure where a child's cp1252 em dash blew up inside subprocess's reader THREAD, returned `stdout=None`, and surfaced as `TypeError: NoneType + str`; and the project-key test stops pinning POSIX, since a canonical Windows path legitimately carries a drive letter. 1,556 tests. **Third pass, on the first COMPLETE Windows run this project has ever had.** With the Ctrl+C bug gone the leg finally executed all 1,558 tests instead of dying at 645, and the honest residue was 23 — not the 6 the truncated run had implied. Most were one shape: paths leaking a host separator into artifacts that are compared. `shlex.split` in POSIX mode EATS backslashes, so the screener saw `python3 D:\\a\\repo\\taskplane\\tp.py` as the single token `D:arepotaskplanetp.py` — every path-shaped screen on Windows (the tp.py read-only exemption, the deny globs) was matching a mangled string. Tokenization now treats a backslash as a separator there, which fails toward MORE recognition, the safe direction for a deny screen. Dispatch briefs are cross-host artifacts whose parity goldens are compared byte for byte between Claude and Codex, so `role_instructions` and the worktree path are normalized before they enter one. The rest were POSIX assumptions in the suite itself, each of which had been quietly reporting a Windows-only failure as a product defect: the DoD test command was `echo ran >> log; exit 1`, which cmd.exe neither chains nor exits from, so every failure-caching case asserted against a run that never happened; `multiprocessing.get_context("fork")` does not exist there; the lock-fallback cases patch `fcntl`, which is POSIX-only, and now skip with a reason rather than erroring. 1,556 tests. **Fourth pass, on the complete failure list.** The eleven cases the first paste had truncated included four more product defects, each one a place where taskplane's own output differed by host. **Artifact bytes:** `json.dump` through a text-mode file turns every `\n` into `\r\n` on Windows, so the SAME state written on two hosts produced different bytes — and taskplane FINGERPRINTS these artifacts and compares them byte for byte, which is exactly how the audit differential caught it. Writes are newline-neutral now. **The regression gate:** `approved_test_roots` returned `{'taskplane\\tests'}` while the radius held `taskplane/tests/...`, so the containment check that keeps fallback discovery inside approved roots matched nothing — a gate that silently stopped constraining. **A symlinked-workspace false positive:** the guard refusing a symlinked test root also refused the workspace itself when reached through an alias, because POSIX resolves a trailing `/.` and Windows does not; a supported spelling became a hard error. **The bare-root guard:** `os.path.expanduser` consults USERPROFILE on Windows and never HOME, so the check that refuses to scope a contract at the session home — the one written after the locked-contract incident — was inert on a host that names its home the other way. It now protects every home the environment names, which can only add protected roots. **Fifth pass — and a scoping rule for the skips.** Two of the remaining failures were regressions the fourth pass caused: normalizing `role_instructions` into `/` shape (correct — briefs are cross-host artifacts) left three golden scrubbers unable to match a host-shaped `PLUGIN_ROOT`, so on Windows the scrub silently did nothing and the compare failed against a real absolute path; and the regression-gate expectations still asserted `os.path.join(...)`, the HOST's shape, against a radius that is `/`-shaped by contract — which is the very property that makes the containment check work on every host. Both fixed at the assertion, with that contract written into the module docstring so the next reader does not have to re-derive it. The genuinely POSIX-only cases skip rather than being ported, and the skip is now scoped to the METHOD that cannot run instead of the class around it: `TestFileLock` was skipped whole on the strength of three cases that patch `fcntl.flock`, which also silenced three lifecycle cases that patch nothing — and on Windows, where there is no fcntl, the mkdir path is not a fallback at all but the production lock, so the broad skip had left the lock taskplane actually uses on that host completely unexercised. Every remaining skip states why it cannot run there, and none of them stands in for a missing product test. 1,558 tests, three more of which now execute on a host without fcntl. **Also: one skill was renaming itself in the host UI.** The Claude skill picker showed nine slugs and one spaced, mixed-case product name — `taskPlane Design` — because `skills/tp-design/agents/openai.yaml` carries a `display_name`, and that Codex interface file also ships inside the Claude package, so a label written for one host renamed the skill on the other. A skill has up to three names (the directory, the frontmatter `name:`, and a host `display_name`) and nothing tied them together. All three are now the slug — which is what the user types — the ten SKILL.md headings share one `# /<slug> — <tagline>` shape, and a gate in the release-freshness suite asserts all three surfaces against the directory so the next interface file cannot rename its skill silently. Verified by reintroducing the exact defect and watching the gate fail. **Sixth pass — three more paths leaking the host separator into a cross-host artifact.** The Windows leg, now finishing the whole suite, failed only the two golden brief compares, and both named the same class the fourth pass had opened and not finished: `role_instructions` was normalized, but the worker workspace (`.tp-work\\t1`), the dashboard pointer (`.taskplane\\dashboard.html`, built with `os.path.join` though it is a logical pointer and never a filesystem operand) and the published artifacts root still rendered host-shaped. Dispatch briefs are compared byte for byte between Claude and Codex, so the same task rendering differently on two hosts IS the defect. All three are normalized on the way OUT — `posix_workspace` sits beside `to_posix` and returns a copy, so stored state keeps the host shape every filesystem use of it wants. Two new checks: the helper driven with Windows-shaped input, and a walk over the frozen goldens asserting no string carries a backslash, which runs on every host and fails if a golden is ever regenerated on Windows — freezing the divergence instead of catching it. The first version of that walk used a path-shaped heuristic and passed against an injected `.taskplane\\dashboard.html`, which contains no forward slash; the rule is now simply that no golden string may contain a backslash, and the injection fails it. loop.py's line ratchet held at 3,003 without being raised: the helper went where it belonged rather than where it was needed. 1,564 tests. **The yield meter — the harness now reports what it RETURNS, not only what it costs.** `ci_loop_cost.py` has pinned spend since R-0012: lenses fired, engine calls mandated, gates passed. Nothing recorded return, so every question about simplifying the harness — can these six lenses go? — had to be settled by taste, and the only visible number was the one going up. `taskplane/yield_meter.py` records the other half into an append-only ledger in the project store (so it accumulates across checkouts and months, unlike rotated runtime state). Two measurements and only two: **lens yield** (reviews routed into, findings produced, what became of them) and **escape** (which gate caught each finding — in-task is cheap, at-review is expensive, after-sign-off is the one that reached a user). What is a fact and what is a guess stay apart: `caught_at` is read from loop state, the stage that INTRODUCED a defect is not knowable here so it is neither stored nor reported, human verdicts (`tp yield mark <fp> acted|dismissed`) are never blended with the weak inference of 'stopped recurring', and a finding nobody has judged yet reports as unknown rather than being rounded to zero — which would slander a lens — or to acted, which would flatter one. The evaluate verdict reports per-lens blocker COUNTS rather than identities, so those shape the escape picture and can never be dispositioned; manufacturing fingerprints for them would be the fabricated precision the meter exists to replace. Explicit disposition is asked for on BLOCKERS only — the findings a human already reads at the gate. **The instrument gates nothing**: it cannot fail CI, block a loop or change a verdict, the engine never reads its ledger (pinned by a test), every write is best-effort, and deleting the file leaves taskplane unchanged — the property that makes it safe to add and easy to remove if it does not earn its keep. Engine-side footprint is an import and one call at the single point every gate transition passes; loop.py's line pin was raised 3,003 → 3,005 on the record rather than hiding the hook somewhere it did not belong. Sixteen tests, and each one was mutated and observed FAILING before it was kept — which caught a decorative one: the garbage-input case passed with the outer exception guard removed, because the inner readers were already total, so it proved them and not the net. 1,580 tests. |
| **v2.7.3** | **The untested trigger was inert on the routing path that actually runs.** v2.7.0 gave `qa` an `untested_trigger` so a change shipping no tests could reach the lens whose Blocker is exactly that. It was verified only through `lens.route(files)` with no stage — the LEGACY path — where it appends reason TEXT after applicability has already been decided. On the stage-aware path (`route_verdicts`, which is what every real review uses) it changed nothing: `qa` stayed `n/a` on precisely the change it exists for, because the qa detector scores test constructs found in the diff and a diff with no tests has none to score. Same shape as v2.6.0's A4 fingerprint defect — a guardrail verified through the path it does not govern — and found by an independent Codex review, not by this repo's own tests. Absence is now applicability evidence IN the engine: an `untested_trigger` lens earns normal path-signal weight when the change carries code and no test path, so it routes light without manufacturing depth, and the stage profile still governs (during `build`, tests legitimately may not exist yet). Second defect, same review: test detection matched SUBSTRINGS, so `contest.py`, `latest.py`, `specification.py` and `protest/` all counted as tests and suppressed the trigger on real code changes. Detection moves to a new dependency-free `taskplane/path_roles.py` that both routers import — segments and filename patterns (`tests/`, `test_*.py`, `*_test.go`, `*.test.tsx`, `*.spec.ts`, `FooTest.java`, `conftest.py`), never substrings — so the legacy reason text and the engine verdict cannot drift apart again. The qa negative fixture was an argparse CLI, a clean negative only while the trigger was inert; it is now a non-code tree that records why it must stay one. Eleven regression tests, every routing one passing `stage=` explicitly, which is exactly what the original verification omitted. The frozen evaluate-stage golden was regenerated: at `build` the profile still excludes qa, but the engine now reports its verdict honestly as light rather than n/a. No pin moved: review cost holds at 7.75 mean / 19 max fired (pins 8.0 / 20). The Claude plugin package also gains a scripted, deterministic build (`scripts/package_claude.py`, `--ext plugin` for file-upload installs, CI-gated for validity and byte-identical rebuilds) — it had been hand-assembled every release, so its membership was whoever built it and no gate could catch a missing member; the OpenAI archive has had one since 2.5.0. 1,541 tests. |
| **v2.7.1** | **Codex CI patch.** The v2.7.0 workflow placed `${{ runner.temp }}` expressions in job-level environment values, where GitHub validates the workflow before the `runner` context exists. Both branch and tag runs therefore failed before creating any jobs. The Codex-host markers and isolated state paths now live on the native-host pytest step, where `runner.temp` is valid, and a CI lint regression test pins that evaluation boundary. No runtime behavior, guardrail, or package membership changed; 1,530 tests. |
| **v2.7.0** | **Lenses 2.0 — all 26 review lenses rewritten against current industry practice, and the review fan-out put on a budget.** Each lens was reviewed individually rather than in batches, against credible dated sources whose authority AND limitation are recorded; where the best available source was a vendor blog it is downgraded and said so, and four superseded citations were caught and corrected (an OWASP edition, a retired web-vitals metric, a stale database version, a platform API level). **The biggest fix was routing, not content:** twelve of twenty-six lenses could not fire on the change class they exist to judge, so their declared Blockers were unreachable by construction — the security lens could not see `.yml`, `.lock` or `Dockerfile`, so a compromised CI workflow never reached the lens that owns supply-chain risk; the i18n lens fired only on translation files, so hard-coded strings in a component never triggered it; `qa` had no baseline, so a change shipping with no tests at all could not reach the lens whose Blocker is exactly that. All twelve closed. **Structure:** the routing rows and review guides move out of the generators into `lenses/_catalog_data.json` and `_guides_data.json` — at 26 lenses x ~7 multi-line checks a Python literal had stopped being editable — and the routing data is now VALIDATED at generation time, so a task type outside the vocabulary fails loudly instead of becoming a silent dead key (eleven such keys were proposed during the work and caught). The generator gained a per-lens preamble, a standing caveat after the checks, a routing note, and a fix for a multi-clause boundary that rendered as "anything under <all of it> belongs to that lens" — naming no lens at all. **Cost:** `qa` fires on a change that adds no test file rather than on every code change (measured: 2 of 40 real changes versus 32, same defect reachable; `TASKPLANE_QA_BASELINE=1` restores baseline), and `scripts/ci_loop_cost.py` now pins the review routing surface over a frozen 20-shape corpus alongside per-task cost, with a deep-cap invariant that fails if the router's budget ever stops being applied. **Same-version release hardening:** the regression gate is fail-closed, discovers non-Python test runners, scopes each task's baseline comparison to its own graph-aware change set, and requires acceptance ownership; Codex briefs now carry collision-safe native task identities, role binding, model and reasoning effort across serial, parallel and lens dispatch, with strict retry-safe verification plus bounded advisory lifecycle hooks; the OpenAI archive ships README, CHANGELOG and complete referenced docs, while a required Codex-host CI leg and archive-membership tests pin the installed experience; legacy import/review scratch trees were removed and ignored. No manifest version changed. 1,529 tests. |
| **v2.6.0** | **v3 Phase 3 plus the performance and review-honesty work the human gated ahead of sign-off** (10 tasks, 5 parallel waves, 1 recorded plan amendment, human-gated plan and sign-off). **Phase 3 (R-0007..R-0011):** every engine-correctness gap the loop hit while governing its own build is closed — `loop claim` refuses in serial mode, per-task DoD excludes loop-owned artifacts, the DoD test subprocess sanitizes the wave slot, the router audit surfaces unattributed findings instead of dropping them, stage emission fails open to the mandatory Task path, and components.yaml floors clamp. Routing precision: the refinement forecast is recalibrated against a pinned corpus, the planner mechanically orders brief-shape changes before golden regeneration, symbol-less big files fold into `::core` honestly, component lens maps union live requirement keywords, and the fixture discount gains a product-directory guard. Host portability: Windows slot activation, `components.yaml` in the default deny family. Docs currency: all five stale skills truthed up, harness rules single-sourced, the freshness gate extended to tp-go, and the 34-flag exemption burned to empty behind a generated `tp help --md` CLI reference. **R-0012 performance:** a month of dogfooding had grown per-task cost about thirteenfold — the product of four independent growths nobody measured together. The DoD test command is now CITED rather than re-run when the same command already completed over byte-identical content under the same engine and env (about six suite executions per agent became one per tree state, and every citation is a traced fact where "I ran the tests" was narration no gate could verify); `tp loop evidence` hands an evaluator the suite result, diff, criteria, routed lenses and graph obligations in ONE call with judgment slots empty, and a bundle submitted unchanged is REFUSED; and `scripts/ci_loop_cost.py` pins what the engine mandates per task, so governance finally carries a budget like everything else. **R-0013 review honesty:** A4's validator-fingerprint refusal had shipped completely INERT — it stamped the submitting process, and every skill resolves the CLI through one installed plugin root, so producer and validator were always the same build; submissions now also stamp the engine in the workspace the evidence came from, which is what a wave worker actually ran. A finding may block a gate only if it carries a claim (trigger, outcome, repro), so commentary stops rendering as a bug — of this review's own 21 findings, roughly seven were not defects. `"set up taskplane"`, the phrase the session-start hook tells users to say, routed to nothing after a docs task narrowed tp-go's triggers; restored on the facade. And the suite's `mkdtemp` calls leaked 185,541 temp workspaces (about 30 GB) until the container's disk filled — fixed at the root, under both runners. No gate was loosened anywhere; two got stricter. 1,415 tests. |
| **v2.5.1** | **Post-release review follow-ups.** Dedicated Codex-compatibility review of Phase 2 against the SHIPPED package under a simulated Codex host: byte-identical dispatch verified, full suite green with zero host leakage, workflow files correctly absent — one real gap found and fixed: `tp onboard` on a Codex host now prints the Codex plugin-tooling install path instead of the Claude org-admin universe (regression-tested). README rebuilt lean: 659 → 320 lines with a new "What taskplane does" feature-definition section; onboarding detail, specialist routes, and Claude Tag moved to docs/{onboarding,specialist-routes,claude-tag}.md — every content pin kept, zero test edits. The tp-help tour now covers routing v2, decomposition, waves, audit cadence, the regression gate, and the install-truth pointer (it had been stale since v2.2.0). 1119 tests. |
| **v2.5.0** | **v3 Phase 2 — all four streams, built and reviewed by taskplane's own governed loop** (9 tasks, 4 parallel waves, 1 honest fix cycle, human-gated design/plan/sign-off). **R-0003 graph decomposition:** `tp graph scan --decompose` derives a `components` layer (directory convention + import cohesion + AST clustering for ≥600-line files; floors 8 files / 600 lines / 2-file·4-symbol·120-line clusters, overridable via components.yaml) with per-component lens maps; reviews route the capped union of touched components' lenses, re-evidenced on the live diff with `component_attribution`, under a fail-open ladder (component → module → breadth=all) whose every rung only WIDENS — structurally tested, and cache invalidation covers graph drift (`graph_sig`). **R-0004 governed flows:** execute/evaluate/fix waves each compile to ONE Dynamic Workflow run between human gates (`--emit workflow|task|auto` on `loop wave`/`loop next`), with the byte-identical Task fallback as the reference path and the only Codex path — parity goldens extended to all three stages, every gate provably reachable with workflows disabled. **R-0005 install truth + docs:** the README now leads with the member reality (org members cannot install from GitHub — admin marketplace or file-upload; claims live-verified against the Claude docs), per-host quickstarts, `tp onboard` detects install context, and docs/routing-and-flows.md documents every new surface (CI-drift-checked). **R-0006 routed reviews + debt burn-down:** evaluate routes stage=build (em full-catalog untouched), fixture-class signals discounted ×0.25 (i18n/mobile inflation fixed — observably, on this release's own review), audit machinery extracted to `taskplane/audit.py` under a byte-frozen differential, and the routed-audit hybrid MEASURED at −0.52%/3 escaped findings → honestly DECLINED per the pre-declared bar (decision 0012). EM review: 1 HIGH (`_scan_hash` containment) + 1 MED (cache graph-drift narrowing) fixed in-tree pre-sign-off. 1120 tests. |
| **v2.4.0** | **Intelligent lens routing + workflow review wave — built BY taskplane itself** (6 governed tasks, 4 parallel waves, design + plan + sign-off human-gated, 0 fix cycles). **R-0001 routing v2:** stage profiles in `catalog.json` (design 8 · build 5 · review 26) plus a signal engine (`lens_signals.py`) that scores each lens against the actual diff — paths, content, density, graph — and returns `deep` / `light` / `n/a`, where n/a always carries machine-checkable negative evidence (e.g. "0 i18n markers"); cap-8 budget demotes (never drops), security floors on enforcement diffs, engine failure fails OPEN to breadth=all, and the legacy routing surface stays byte-identical. An audit sweep every Nth review (`TASKPLANE_AUDIT_EVERY`) diffs findings against the routing and auto-files any n/a-lens finding as a router regression that blocks sign-off. **R-0002 review wave:** `workflows/review-wave.js` runs the whole lens fan-out as one Claude Dynamic Workflow (schema-pinned findings), with a MANDATORY byte-identical Task-dispatch fallback — Codex behavior unchanged, pinned by CI parity goldens. Dogfooding found 4 real engine bugs, all fixed: worktree store resolution, the scope-precedence deadlock, its provenance exploit (literal scope overrides now require the human-approved plan's `plan_minted` mark — CLI `--scope` never), and the sign-off aggregate flagging loop-owned artifacts. Hardening from the EM review: root-level `.env`/`secrets/` join the sacred deny family, detector reads are realpath-contained, the workflows kill-switch accepts `false/no/off`, and bare `n/a` coverage without evidence blocks the em gate. 954 tests. |
| **v2.3.1** | **Graph-scoped regression gate + review discipline** — taskplane now verifies each change for *actual regressions within its blast radius, right away at the DoD gate*, and stops reviews from reporting taste as blockers. The gate (`taskplane/regression.py`, opt-in via `dod.regression_gate`) has two tiers: **Tier 1** selects the tests covering the changed modules (via the dependency graph + a test-import index), runs them at the change's baseline commit in a throwaway worktree and again now, and blocks only on a test that was **green before and fails now** — a real regression, named; a pre-existing failure never blocks. **Tier 2** flags an enforcement/public entry point (screen/DoD/hooks/CLI/CI invocation) changed with **no covering test** — the class of regression a test-diff can't see, which is exactly how v2.3.0's CI-invocation break shipped. Reviews now classify every finding `regression | pre-existing | observation`: only a regression, or an unclassified high in the change's own diff, gates sign-off (`loop.finding_blocks`/`classify_findings`); pre-existing debt and taste are surfaced and tracked, never block — so a 26-lens sweep reads as "N block · M to triage" instead of "100 issues". The gate degrades visibly (says so) when a baseline or graph is unavailable, and never crashes the DoD it guards. 768 tests. See docs/regression-gate-design.md. |
| **v2.3.0** | **Fix all 122 findings from the whole-codebase 26-lens review of v2.2.1** (12 high · 57 med · 53 low), under a binding rule: **no fix may reduce a guardrail** — every enforcement change is strict-or-stricter, proven by a 66-case differential battery (zero block→pass regressions vs v2.2.1) and a `TestNoLoosening` suite. Highs: the single active-contract slot is replaced by per-task contract files (`TASKPLANE_TASK`) that fail closed to the **most-restrictive union** when ambiguous, so parallel lens/wave agents can no longer overwrite each other's governance; Windows hooks fail **closed** when the plugin root is unset (was fail-open); the read-only screen now blocks `python file.py` / `sh script.sh` / `bash -lc` interpreter escapes (was `-c`-only); the test suite is isolated for **both** runners (`python -m unittest` no longer writes to the developer's real `~/.taskplane`); the EM gate maps unknown/`blocker`/`major`/`critical` severities **up** to high (a mis-labelled blocker can no longer slip sign-off); `tpFire` no longer fakes success in a static view; paged renders enforce the ≤14 KB budget in real bytes; a malformed trace event renders degraded-but-visible instead of crashing; budget exhaustion is disclosed in the headline. Durability: one atomic-write + never-silently-lock-free primitive (`atomic_write_json`/`load_json`/`file_lock`, mkdir fallback on FUSE/Windows) now backs graph.json, mode.json (fails toward **private**), tracks.json, meter.json, the requirements index and loop state — a torn write or corrupt control file fails closed with a remedy instead of silently losing data. Plus single-source versioning (`tp version --verify`, CI-gated), `.gitignore`-aware graph scans, incremental kb-lint, honest Go-scanner limits, pragmatic i18n/RTL/mobile fixes, `tp gc` for runtime artifacts, and doc/CI truth-up (authority matrix, state-spec, configuration reference, lens-catalog drift gate). 744 tests. |
| **v2.2.1** | **Fix all 32 findings from the full 26-lens self-review of v2.2.0** (5 high · 11 med · 16 low, every severity addressed) plus 2 renderer-contract findings the human filed mid-review. Highs: worker submissions can no longer be misattributed across tasks (`--task` is validated everywhere), gate transitions apply under the state lock so a parallel wave worker's update is never clobbered, the design graph baseline re-captures after a legitimate rescan instead of deadlocking, the pm gate is fail-closed (an authored requirement must exist before Define advances), and post-approval design tampering is now pinned by tests at all four gates. Structure: Design Contract validation extracted to `design_contract.py` (loop.py −18%), policy/contract-id normalization unified in depgraph, DoR checks are pure (only the gate applies mutations). Renderer: wave-board lanes derive status from findings files so re-rendering shows the live fan-out, and paged dashboards mandate byte-for-byte verbatim rendering. Java packages no longer collapse across group ids; fingerprints never silently disable (HEAD fallback); anonymous approvals are recorded as `(unattributed)` with a warning. 477 tests. |
| **v2.2.0** | **First-class Design before Build** — `taskplane design` is now the proposed-HOW entry point, distinct from Product's WHAT, Build's realization, and Review's judgment. Design is read-only toward product code and produces a human-approved, fingerprinted `taskplane.design/v1` contract with alternatives and trade-offs, selected approach, current-state grounding, proposed modules and dependency edges, named contracts, bounded graph depth, graph DoR/DoD, exact acceptance-to-validation mapping, risks, failure modes, observability, rollout/rollback, and a conditional technical visual. Complex Build work can opt into the Design phase; Design-only work ends after approval. Plan must cover the approved contract, Evaluate rejects stale evidence, and Engineering Review must prove module/edge/contract conformance with no unexplained drift. The new solution-design lens is intentionally distinct from architecture, trade-offs, and UI/UX design, bringing the catalog to 26. Claude and Codex use the same engine, skill, role, and human gate. 456 tests. |
| **v2.1.0** | **AI software delivery with proof, not agent self-reporting** — taskplane is now positioned and exposed as a simple AI software-delivery control plane: `taskplane build`, `taskplane review`, and `taskplane status` route into the existing specialist system while the full harness remains mandatory. Workers submit source-and-artifact fingerprints but cannot advance lifecycle state; the orchestrator independently gates and rejects missing, stale, out-of-scope, under-tested, or under-reviewed work. Requirements carry dependencies and named contracts into planning; graph-aware DoR refreshes before approval, requires every new module to be declared, and bounds distributed traversal at explicit contract/resource nodes. Graph-aware DoD checks realized modules, affected consumers and requirements, engineering narrative evidence, the approved aggregate impact policy, and a current graph fingerprint. Workspace identity is canonical across macOS path aliases, rejected plans retain their active contract, evaluator and EM artifacts are fingerprint-bound, and shared agent definitions inherit the host model across Claude and Codex. 444 tests. |
| **v2.0.0** | **One governed delivery plane for Claude and Codex** — the same enforced Definition of Ready, scoped task contracts, evidence-backed Definition of Done, human gates, and full 25-lens review now run across both hosts. Codex compatibility covers `apply_patch`, agent dispatch, host-portable role commands, model inheritance, onboarding, and dashboard fallbacks without bypassing Codex's sandbox or approval flow. Every gate publishes durable progress artifacts so a new session or teammate can resume from the approved plan, dashboard, findings, graph, retro, and `HEADLINES.md`. The dependency graph now actively shapes delivery: hub changes escalate architecture review, execute/fix briefs receive blast radius before editing, lens agents receive impact context, prior snapshots surface during onboarding, and reviews capture SQL, HTTP, and messaging edges the import scanner cannot see. Import-node unification and module-id impact fixes restore accurate dependent counts. 430 tests. |
| **v1.6.0** | **Codex support — same governed loop, second host** — taskplane now packages as an OpenAI Codex plugin from the same repo: the shared hook screens Codex `apply_patch` (every patch target, move destinations included; opaque patches fail closed) and `Agent` dispatches, contracts written on Claude stay valid via tool aliasing, and allowed actions preserve Codex's own sandbox/approval flow. Governance hardened for both hosts: DoR now refuses to start unready steps, PASS advances only on evidence (task tests + evaluator criteria/lens verdicts + the full-catalog review), and a failed final DoD blocks sign-off. Dashboards render inline widgets where the host supports them and fall back to `HEADLINE:` + a generated HTML artifact; Codex inherits its session model for every tier unless explicitly mapped. Claude behavior unchanged. 409 tests. |
| **v1.5.4** | **Lens coverage & dependency graph, surfaced where you decide** — new review lenses kept drifting out of the visualization and the dependency graph kept getting skipped. Both now derive from their single source of truth and appear automatically: the dashboard renders a **lens-coverage panel from `catalog.json`** (add a lens, it shows up — no hand-editing), the findings dashboard carries a **blast-radius graph panel** (an empty graph on a polyglot repo *explains* the cross-service gap and how to record edges with `tp graph edge`, instead of showing nothing), and the never-skippable `HEADLINE:` now reports lens coverage (`N deep / M sweep of 25`) and modules touched. The same render flow — HEADLINE first, `--paged` widgets, catalog in the context tab, graph in the graph tab — is now shared by **every** command: go, engineering, product, and north-star. Rendering is pinned to the cheap/standard model tier. |
| **v1.5.3** | **Render-reliability contract for inline dashboards** — the decision data in a dashboard no longer depends on one big widget that might get skipped. Every dashboard command prints a never-skippable plain-text `HEADLINE:` with the key numbers; `tp findings --paged` / `tp dashboard --paged` split a large view into ordered, self-contained fragments (summary → high → medium → low), each under 14 KB, rendered one after another; and the skills now mandate rendering every page in order rather than summarizing. Small reviews still render as a single widget. |
| **v1.5.2** | **Hardening from the full 25-lens self-review of v1.5.1** (62 findings across every severity, all addressed): write-screen now covers `git apply/am`, `patch`, `sort -o`, and `cp/mv -t` and treats unscopeable mutation as default-deny; dependency-graph HTML escapes repo-supplied names (XSS); shared-store decision ids are collision-free; `publish`/`set_status`/`supersede` and the dispatch queue are lock-serialized; the team store is committable (anchored gitignore); private mode refuses inside Tag; `init --plan` scaffolds into the right store; docs (README, PRIVACY, state-spec, loop-design) truthed-up to the plan-aware store; keyboard + ARIA on the dashboard; and a raft of smaller correctness/UX fixes. 373 tests. |
| **v1.5.1** | **Hardening from a full 3-lens review of v1.5.0** (22 findings, all fixed): shared-store bookkeeping no longer trips the DoD gates; private mode refuses loudly inside Claude Tag instead of silently writing to the shared store; `init --plan` scaffolds into the right store; loop coordination state stays per-user (share knowledge, not the state machine); publish hardened — corrupt shared index aborts instead of erasing history, stale/lost markers repair via content-based idempotency, collision-free hash-suffixed shared ids, flows now push too, unknown ids and malformed entries reported; KB writes are atomic + locked; repo-supplied shared config carries a visible notice; plan/privacy settings follow the repo across checkouts. |
| **v1.5.0** | **Plan-aware onboarding, private mode & share push** — onboarding asks whether you're on a personal or Team/Enterprise plan (changeable any time with `tp share plan`): personal keeps knowledge in the private store (`~/.taskplane`), team/enterprise shares it in-repo (`.taskplane-kb/`, Claude-Tag compatible, inherited by every clone). On a team plan you can still work privately (`tp share set private`) and publish selected decisions to the team later with `tp share push [--ids …]` — deliberate and idempotent, like pushing commits. |
| **v1.4.0** | **Claude Tag support (beta)** — run governed work as your org's @Claude in Slack: `TASKPLANE_STORE=repo` persists the knowledge store inside the repo (Tag's sandbox is ephemeral — the KB travels with the branch/PR), `tp loop approve --by "<who> — '<their words>'"` makes every human-gate pass attributable to a real person in the thread, and the new `tp-tag` skill carries the thread protocol: post the gate, stop, wait for a human reply, attach the dashboard. |
| **v1.3.1** | **Remedy required on every finding** — the shared verdict format across all 25 lenses makes `suggestion` mandatory on every blocker/major/minor: criticism must come with a concrete alternative or solution, preferring capabilities your stack already runs. |
| **v1.3.0** | **Current-state grounding** — `tp init` scaffolds `context/current-state.md` (the as-built inventory); once filled it is injected into every task brief, and the design lenses review any design as a *delta against what exists*: reinventing an existing component or contradicting as-built reality is blocker-class. Grounding panel in the dashboard. |
| **v1.2.0** | **Decision registry** — `tp decision`: structured ADRs with lifecycle, alternatives (`option \| gained \| given up`), and supersede chains; accepted decisions linked to a task's modules are always injected into that task's brief. Plus three new design lenses: trade-offs, services selection, time-to-market (catalog: 25). |
| **v1.1.0** | **Dashboard v2** — auto-rendered on every gate and step; clickable loop rail + step journey with per-step execution and decision details; full acceptance criteria and execution plan shown for review; always-on stats with the agent → model table. |
| **v1.0.1** | Dispatch verification (`tp loop verify-dispatch` + opt-in Task-dispatch hook), FUSE-safe store cleanup, richer task statuses (`done`/`external` satisfy dependencies), model-tier setup in onboarding. |
| **v1.0.0** | First public release — enforced task contracts (PreToolUse hook), human-gated Evaluate-Loop, context-routed lens catalog, requirements engine, external per-project knowledge base, model tiers, on-demand north-star review. |

SHA-256: 87c74fd8308ce5ca8a2bbe9e5f1a5a0f3543badc0a3535822fd2b28b37149663