← Files codex-sdlcARCHIVED FILE
docs/workflow-benchmark.md
5.09 KB · Oct 3, 2026 · 06:35 UTC
# Full and Compact workflow benchmark This is a reproducible local framework benchmark, not a measurement of real agent implementation time, human review, model reasoning, or token savings. ## Run it From a complete repository checkout with its dependencies installed: ```sh node scripts/benchmark-workflows.mjs --repetitions 3 ``` The script captures baseline commit `17bde3f88bb3b1633c2eb1a8ea9e5b565449076f`, builds it in an isolated ignored directory, builds the candidate, and compares baseline Full, candidate Full, and candidate Compact. The baseline commit must be available in local Git history. `--baseline-only` records just that baseline. Repetitions are bounded from 1 to 20. Results are written to `build/benchmarks/baseline.json` or `build/benchmarks/comparison.json`. The script creates and removes disposable fixtures; it does not use a customer project, actual database, external service, or live agent. It checks the candidate source/assets fingerprint before and after the comparison and refuses a completed result if they changed during measurement. ## Comparable workload Every fixture starts from the same small order-list implementation and uses the same source, acceptance test code, and data seed. Five checks cover authentication, tenant permission, one-based pagination, invalid/empty-page behavior, and safe output mapping. Three deliberately broken variants exercise permission leakage, incorrect page offset, and incorrect mapping. Each must fail its check and be refused by the handoff operation; the corrected implementation must pass. The script performs the prescribed runtime lifecycle with clearly synthetic role and review records. Full uses fixed documentary fixture content; Compact uses a canonical specification and QC input with generated views. No real PM, BA, implementation agent, or independent human reviewer is simulated as a measured person. Final human acceptance remains pending in every successful fixture. Candidate Full must retain the same graph as baseline Full. Source/test hashes and broken-variant outcomes must agree across all suites. Compact combines integration and QC, so repeated invocations of the same acceptance suite differ by design even though the acceptance criteria and test source are identical. ## What the report means | Metric | Meaning | | --- | --- | | Task count | Tasks required by the selected graph for this web-only fixture | | Required outputs | Distinct required workflow output paths | | Fixture-authored required outputs | Required files populated from fixed fixture content, not measured human/LLM writing | | Generated required outputs | Required files generated by the runtime | | Semantic inputs | Separate request, assessment, specification, controls, outcomes, and QC input categories | | Auxiliary generated artifacts | Assignment, receipt, changed-file inventory, and collector JSON/stdout/stderr files | | Runtime operation time | Scripted lifecycle/helper work, including validation and I/O | | Collector wrapper time | Process execution plus evidence publication overhead | | Collector command time | Duration recorded inside collector evidence; excludes wrappers and agent work | | Fixture wall time | Entire scripted fixture, including setup and Git/file operations | Task and output counts for this fixture are Full **7 tasks / 24 required outputs** and Compact **5 tasks / 11 required outputs**. This demonstrates a smaller prescribed workflow, not an equivalent percentage reduction in real feature delivery time. Compact adds structured inputs and generated views, so total generated file count can increase even while the required documentary workload decreases. Use medians from the completed report, and retain the raw samples. A few local repetitions are sensitive to caches, system load, and process startup. They do not establish a production speedup or justify applying the percentage to a previous multi-hour agent run. A future real-agent pilot should use matched starting code, requirements, environment, model policy, and independent acceptance evaluation, with blocked and failed runs retained in the results. ## Recorded development result The completed comparison on Node.js 24.19.0, macOS arm64 used three repetitions per suite. Canonical fact/claim mapping, implementation/test source hashes, all three defect outcomes, and pending human acceptance were checked across all nine runs. | Metric | Baseline Full (`17bde3f`) | Candidate Full | Compact | | --- | ---: | ---: | ---: | | Tasks | 7 | 7 | 5 | | Distinct required outputs | 24 | 24 | 11 | | Runtime operation calls | 42 | 42 | 37 | | Median scripted fixture wall time | 4.55s | 4.43s | 3.44s | | Total generated required + auxiliary artifacts | 27 | 27 | 29 | Compact's fixture wall time was about 22% lower than candidate Full in this experiment. The reduction in required outputs does not mean every file count or all writing decreased: Compact adds structured semantic inputs and generated views. The small Full-to-Full timing difference is not treated as a demonstrated optimization. Earlier incomplete measurements with noncanonical Full claim fixtures are superseded by the corrected completed comparison.
SHA-256: 533c89e2c15641472926436e986002ab5c6bfaa5f8ac04ce6725621e630bf5b4