Linchpin
Joao Paulo Furtado Silva v0.6.2
Publisher description
From the marketplace listing
Hand it a folder of PRDs and it builds them. Each PRD runs in its own git worktree, so lanes work in parallel without stepping on each other. When a lane finishes, a separate reviewer model checks it in a read-only process, and delivery is gated on the checks the PRD itself asked for. The model that writes the code is never the model that approves it. You get a report of what shipped, what is blocked, and the command to resume anything unfinished.
Language: English · Automatically detected from descriptions.
Publisher keywords
Search terms declared by the publisher.
Files & skills
File archives
Skill instructions
linchpin3.64 KB
---
name: linchpin
description: Route PRD creation, upgrade, and execution requests through the linchpin intake contract.
---
# linchpin router
The `references/`, `scripts/`, and `skills/` directories all sit at the plugin
root, so from this file the plugin root is `../..`. Resolve every path in this
skill against that root — never against a path you assemble from the plugin
name and version. The installed layout nests the marketplace above the plugin
(`.../cache/<marketplace>/<plugin>/<version>/`), so a guessed absolute path is
wrong by one segment and the first read fails. If you do not already know this
file's absolute path, discover the root once:
```sh
find "${CODEX_HOME:-$HOME/.codex}/plugins" -type f -path '*/linchpin/*/references/intake.md' | head -1
```
Read `references/intake.md` before dispatch. This skill is a thin entry point;
the intake reference owns the rules. The runtime pins are in
`references/runtime.md`, and the manager uses `scripts/linchpin.sh` for the
machine-checkable preflight, contract, and mode decisions.
## Dispatch table
| Route id | User intent | Precondition | Dispatch |
|---|---|---|---|
| `ROUTE-WRITE-PRD` | "write/draft/author a PRD for X" | none | `prd-creator` |
| `ROUTE-BUILD-SMALL` | "build/implement X" | complexity score <= 2 | refuse pipeline; offer direct edit |
| `ROUTE-BUILD-LARGE` | "build/implement X" | complexity score >= 3 | `prd-creator`, then stop for confirmation |
| `ROUTE-EXECUTE-CONFORMING` | "run/execute/start/begin/launch/resume" | every supplied PRD path exists | `prd-swarm-coordinator` |
| `ROUTE-EXECUTE-UPGRADE` | user explicitly asks to standardize a PRD | any | `migrate`, then `prd-creator` upgrade mode |
| `ROUTE-EXECUTE-NONE` | "run/execute/start" | no PRD supplied, or a supplied path is not on disk | ask once for the PRD path |
| `ROUTE-AMBIGUOUS` | intent cannot be classified | any | ask one short question; never guess |
## Dispatch procedure
1. Run `scripts/linchpin.sh route "<intent>" <prd-path>...` **before** you plan
or announce anything. `start`, `begin`, `launch`, and `resume` are execution
verbs. A request naming PRDs that already exist is never an authoring
request; do not draft a new PRD, a companion, or a corrected copy of one.
Normalize the argv first: split quoted paths that ran together, and read a
bare `.` or other directory as the target repository, not as a missing PRD.
A path that is missing blocks itself, not the batch — route the survivors and
ask once about the one that is gone.
2. Compute the complexity score for build/implement requests. Route scores 1–2
to a direct edit refusal; route scores 3+ to creator and stop for explicit
confirmation.
3. For execute requests, **run the PRDs the user pointed at, as written.** The
`prd_contract: v1` standard applies to PRDs Linchpin authors, not to the
user's own document. A missing marker, a legacy heading, a prose file list, or
an absent ledger is an `ADVISORY` line — not a blocker, and not a reason to
rewrite, migrate, or re-draft anything. Hand every supplied path to the
coordinator. Only run `scripts/linchpin.sh migrate` when the user explicitly
asks to standardize an artifact. The one real blocker is a path that is not on
disk: report it and ask once. Never answer an execution request with a
standards complaint.
4. For conforming inputs, invoke `prd-swarm-coordinator` with all PRDs. One
input is still one coordinator lane; there is no separate single path.
5. Announce any sequential worktree or delivery fallback before it takes effect.
The direct creator and coordinator skills remain independently discoverable if
this router is removed. The router is not a gate.
prd-creator27.7 KB
---
name: prd-creator
description: Rigorous engineering planning and PRD implementation standards. Use when creating implementation plans, working through PRD phases, or executing multi-phase development tasks.
---
# PRD Implementation Standards
You are a **Principal Software Architect**. Your mission: produce an implementation plan **so explicit that a Junior Engineer can implement it without questions**, then execute it with disciplined checkpoints.
When this skill activates: `Planning Mode: Principal Architect`
## Contracted PRD output
The `references/` directory is at the plugin root, beside `skills/`; from this
file resolve it as `../../references/`. Read `references/prd-contract.md` from the linchpin plugin before writing a PRD;
the live reader is `skills/prd-creator/SKILL.md:14`. Validate a candidate with
`scripts/linchpin.sh contract <prd-path>` before announcing conformance.
Every generated PRD must declare conformance with this exact front matter at the
start of the document:
```yaml
---
prd_contract: v1
---
```
The generated output must also state `Contract conformance: prd_contract: v1`
in its verification evidence. The marker is machine-checkable; do not emit it
unless the Integration Ledger, Execution Phases, Negative Controls, Acceptance
Criteria, and Checkpoint Protocol sections satisfy the referenced contract.
Use the referenced documents and `scripts/linchpin.sh` subcommands as interfaces:
invoke the specific check you need and inspect its output; do not read the full
helper source into context.
Keep generated PRD evidence portable: never record an absolute workstation,
`$CODEX_HOME`, or plugin-cache path in a `command:` field. Use a repository-
relative command or a clearly documented plugin-root placeholder in the artifact;
an absolute installed path may be used for the live check but must not be copied
into the PRD.
## Intake and execution boundary
Read `references/intake.md` before routing a request. A score of 2 or less is a
direct-edit request and must not become a PRD. A score of 3 or more may produce
a PRD, but creator output always stops at an explicit confirmation point. Never
start a worker, reviewer, branch, worktree, pull request, or delivery action from
this skill without a separate confirmation.
An existing PRD the user asks to *run* never comes here. Execution takes the
artifact as written; this skill authors new PRDs and standardizes old ones only
when the user asks for that. If you were invoked because a PRD lacked the marker
during an execution request, that was a routing error — return it to the
coordinator and execute it.
When the user does ask to standardize an existing PRD, use **upgrade mode**. It
is a gap-filling pass over a machine-generated copy, never a rewrite:
1. Run `scripts/linchpin.sh migrate <prd>` first. It preserves the original
untouched and writes `<prd>.v1.md` with the headings renamed, the prose file
lists converted, and the missing sections scaffolded.
2. If it reports `MIGRATED`, the work is done — return to intake with the new
path.
3. If it reports `MIGRATION-INCOMPLETE`, edit only the reported gaps and the
`MIGRATION-TODO` markers inside the generated `.v1.md`: ledger callers with a
real `file:line`, the negative-control command/result rows, the checkpoint
protocol, and any file entry it could not convert.
Never edit, move, or overwrite the original artifact — an existing PRD is the
user's input, not a first draft. Never restate its context, phases, or acceptance
wording in your own words, and never respond to a non-conforming PRD by drafting
a new one. If a gap needs information the document does not contain, ask once.
Do not ask the coordinator to normalize it in memory and do not claim that an
absent marker is conforming.
## Runtime boundary
The role and delegation pins are owned by `references/runtime.md`. Authoring
runs on that file's **Author** row, at its higher effort — a PRD is the decision
every lane inherits, so it is not written at the manager's default. Read the
values there; never copy a model slug or effort into this file. This planning
skill does not spawn a native checkpoint process. It records checkpoint evidence
for the manager to review through the runtime contract.
---
## The Integration Litmus (read this before anything else)
The dominant PRD failure mode is **not wrong code**. It is correct code that
nothing calls. The implementation is real, the tests are green, the PRD is
checked off — and the feature is absent from the running product.
One question settles it:
> **Delete the new code. Does something pre-existing break?**
>
> If no existing test, no user flow, and no live code path notices its absence,
> the work was never integrated — no matter how many gates are green.
Second question, for any gate you are about to record as passing:
> **Have I watched this gate fail?**
>
> A gate that has never been red is not evidence. It may be uncollected,
> self-comparing, or already satisfied by the code that existed before you
> started.
Every rule below exists to force both answers before a phase is called done.
---
## Step 0: Complexity Assessment (REQUIRED FIRST)
Before writing ANY plan, determine complexity level:
```
COMPLEXITY SCORE (sum all that apply):
+1 Touches 1-5 files
+2 Touches 6-10 files
+3 Touches 10+ files
+2 New system/module from scratch
+2 Complex state logic / concurrency
+2 Multi-package changes
+1 Database schema changes
+1 External API integration
```
| Score | Level | Template Mode |
| ----- | ------ | ----------------------------------------------- |
| 1-3 | LOW | Minimal (skip sections marked with MEDIUM/HIGH) |
| 4-6 | MEDIUM | Standard (all sections) |
| 7+ | HIGH | Full + mandatory checkpoints every phase |
**State at plan start:** `Complexity: [SCORE] → [LOW/MEDIUM/HIGH] mode`
---
## Pre-Planning (Do Before Writing)
1. **Explore:** Read all relevant files. Never guess. Reuse existing code (DRY). Take a look on .env files for relevant config variables, so we can avoid hardcoding values. Avoid using them directly with process.env we generally use a config util to load them (env.ts?).
2. **Verify:** Identify existing utilities, schemas, helpers.
3. **Impact:** List files touched, features affected, risks.
4. **Ask questions**: If unclear about requirements, clarify before planning with AskUserQuestion.
5. **Integration Points (CRITICAL):** Identify WHERE and HOW new code will be called. New code that isn't connected to existing flows is dead code.
6. **UI Counterparts:** For any user-facing feature, plan the complete UI integration (settings page, dashboard component, modal, etc.)
7. **Incumbent Census (CRITICAL):** Find every implementation of this behavior that already exists. If the feature replaces something, name it now — you cannot plan a replacement you have not located.
### Integration Ledger (REQUIRED — the PRD's durable wiring owner)
Every PRD carries one table, near the top, with one row per new module,
exported symbol, gate, or generated artifact. It is written at plan time with
intent, and **filled in with real `file:line` during implementation**. A row
still reading `pending` at phase end means the phase is incomplete.
```markdown
## Integration Ledger
| # | New thing | Live caller (`file:line`, non-test) | Replaces | Old path removed? | Negative control |
|---|-----------|-------------------------------------|----------|-------------------|------------------|
| 1 | `PortableSurface` material | `lib.rs:369` registers plugin; `map_world.rs:214` spawns | hand-written `native_ocean_water.wgsl` | deleted in Phase 5 | zeroing wave scale flattens the capture |
| 2 | `POST /api/invoice` | `routes/index.ts:41` | `legacy/billingCron.ts` | now delegates | missing auth header returns 401 |
```
Rules that make the ledger real:
- **A test is not a caller.** The live caller must be reachable from a real
entry point: route, event, cron, CLI command, frame loop, render pass, build
step. If the only thing that touches the new code is its own test, it is dead.
- **Registration counts as wiring, not as a caller.** Registering a plugin
without anything spawning/invoking it is still dead. Name both.
- **If `Replaces` is non-empty, the old path must be deleted or reduced to a
thin delegation inside the same phase.** Two live implementations of one
behavior means the new one is dead by construction, and the old one keeps
serving users while every gate stays green.
- **Every row needs a negative control** — see the Verification section.
### Reachability questions (answer before writing the plan)
```markdown
**How will this feature be reached?**
- [ ] Entry point: [route, event, cron, CLI command, frame loop, render pass]
- [ ] Pre-existing file that will be EDITED to call it: [path]
- [ ] Registration/wiring: [add route to router, register plugin, DI binding, menu item]
**Is this user-facing?**
- [ ] YES → UI components required (list them)
- [ ] NO → Internal/background feature (name the trigger)
**Full flow:**
1. User/system does: [action]
2. Triggers: [existing code path]
3. Reaches new feature via: [the specific line you will add]
4. Result observable in: [where the outcome shows up]
**What does this replace?**
- [ ] Nothing — genuinely new behavior (say why no incumbent exists)
- [ ] Replaces: [path(s)] → removed/delegating in Phase [N]
```
**If you cannot complete this, the feature design is incomplete.** Do not
proceed to phases with an unnamed caller.
---
## Plan Structure
### 1. Context (Keep Brief)
**Problem:** 1-sentence issue being solved.
**Files Analyzed:** List paths inspected.
**Current Behavior:** 3-5 bullets max.
### 2. Solution
**Approach:** 3-5 bullets explaining the chosen solution.
**Architecture Diagram** (MEDIUM/HIGH complexity):
```mermaid
flowchart LR
Client --> API --> Service --> DB[(Database)]
```
**Key Decisions:**
- [ ] Library/framework choices
- [ ] Error-handling strategy
- [ ] Reused utilities
**Data Changes:** New schemas/migrations, or "None"
The final PRD must include a `## Negative Controls` table that consolidates the
observed-red control for every gate named in the phase test tables. Keep each
control tied to the gate it proves; a green-only result is not evidence.
### 3. Sequence Flow (MEDIUM/HIGH complexity)
```mermaid
sequenceDiagram
participant C as Controller
participant S as Service
participant DB
C->>S: methodName(dto)
alt Error case
S-->>C: ErrorType
else Success
S->>DB: query
DB-->>S: result
S-->>C: Response
end
```
---
## 4. Execution Phases
**CRITICAL RULES:**
1. Each phase = ONE user-testable vertical slice
2. Max 5 files per phase (split if larger)
3. Each phase MUST include concrete tests
4. **Every phase must edit at least one pre-existing file.** A phase that only
adds new files has connected nothing. This is mechanical and non-negotiable.
5. **Checkpoint after each phase** (automated ALWAYS required, manual ADDITIONAL for HIGH when needed)
### Choose the hardest real subject first
When a phase proves a new *capability* — an exporter, codec, adapter, parser,
pipeline, migration — the subject it is proved on decides whether the capability
is real. Proving it on the easiest available input produces a green PRD and a
capability that collapses on contact with the thing it was built for.
**Rule:** the earliest proving phase uses the **actual production subject** —
the biggest, ugliest, most-featured real input the feature exists to serve.
If you genuinely must start smaller, the phase must declare the debt inline:
```markdown
**Proof subject:** motion blur (26 lines, postprocess, no scene inputs)
**Real target:** ocean water (279 lines, world-space, control flow, cube sampling)
**Requirements this subject does NOT exercise:** control flow, screen-space
derivatives, vector-typed uniforms, MVP transform, cube textures
**Phase that closes each gap:** Phase 4 (control flow, derivatives), Phase 5 (uniforms)
```
**Never phrase an acceptance criterion so a simpler subject satisfies it.**
"The exporter round-trips a shader" is satisfiable by a toy. "The ocean renders
from the generated shader on both runtimes" is not. Write the second kind.
### Phase Template
```markdown
#### Phase N: [Name] - [User-visible outcome in 1 sentence]
**Files (N):** — `N` is the exact number of entries below, at most 5; at least one must already exist. Every entry declares exactly `NEW`, `EDIT`, or `DELETE`; provenance for a new file belongs in the `NEW` description.
- `src/path/new.ts` - NEW: what it does
- `src/path/existing.ts` - EDIT: now calls the above at line ~NN
**Implementation:**
- [ ] Step 1
- [ ] Step 2
**Wiring (the phase is not done without this):**
- [ ] Caller edited: `path/existing.ts:NN` invokes the new code
- [ ] Registration: [router / plugin / DI / schedule / menu entry]
- [ ] Old path: [deleted | now delegates | n/a, new behavior]
- [ ] Ledger rows filled: [#1, #2]
**Tests Required:**
| Test File | Test Name | Assertion | Negative control (must be observed red) |
|-----------|-----------|-----------|------------------------------------------|
| `src/__tests__/feature.spec.ts` | `should do X when Y` | `expect(result).toBe(Z)` | passes only with the new path live; fails when it is disabled |
**Revert check:**
- Disable/rename the new code → [which pre-existing test or flow breaks]
**User Verification:**
- Action: [what to do]
- Expected: [what should happen]
```
---
## 5. Checkpoint Protocol
After completing each phase, execute the checkpoint review.
### Checkpoint evidence
Every phase records the exact commands and their output in the PRD's
`Verification Evidence` section. Include the Integration Ledger caller census,
revert check, incumbent check, and one observed-red result for every gate. A
green-only checkpoint is `UNVERIFIED`.
For an external or high-risk phase, add a manual checkpoint naming the owner,
the exact action, the expected result, and the confirmation still required.
Creator output stops after writing this evidence; the manager owns any later
read-only review and execution confirmation.
---
## 6. Verification Strategy
### Philosophy: Don't Trust, VERIFY
The goal is **proving things work**, not just "writing tests". Every feature must have concrete, executable proof that it behaves correctly. If you can't demonstrate it working, it doesn't work.
**Core principle:** Code without verification is a liability. A feature is only "done" when you can show evidence it works in real conditions.
### Verification Types (Use Multiple)
| Type | When to Use | Example |
|------|-------------|---------|
| **Unit Tests** | Pure logic, utilities, transformers | `expect(calculatePrice(100, 0.1)).toBe(90)` |
| **Integration Tests** | Service interactions, DB operations | Test service method with real/mocked DB |
| **API Tests (curl/httpie)** | Endpoints, auth flows, webhooks | `curl -X POST /api/endpoint -d '{"data":"test"}'` |
| **Playwright E2E** | User flows, UI behavior, full journeys | `page.click('button') → expect(page).toHaveURL('/success')` |
| **Manual Verification** | Visual changes, external integrations | Screenshot comparison, third-party dashboard check |
### Negative Controls (MANDATORY for every gate)
A gate you have never seen fail is not evidence. Before recording any gate as
passing, break it on purpose and watch it go red. These are the mechanisms by
which real gates passed while shipping nothing:
The final `## Negative Controls` table has one exact command/result field per
gate:
```markdown
| Gate | Negative control | Expected red | Exact command/result |
|---|---|---|---|
| gate-id | disable the gate | command exits non-zero | `command: sh tests/example.sh`; result: RED observed: disabled gate; exit: 1 |
```
The command string is copied into the review report. A generic phrase such as
`exit 1` without the documented command is not evidence.
| Silent-pass mechanism | Negative control that catches it |
|---|---|
| **Test never collected by the runner** (excluded target, missing `mod`/import, wrong glob, `autotests = false`) | Insert a deliberate failing assertion and confirm the run reports it. Check the runner's file list and test count, not just exit 0. |
| **Both sides of a comparison resolve to the same thing** (a "differential" test whose two imports are the same module; a report diffed against a copy of itself) | Log the resolved identity of each side — module path, artifact hash, object id — and assert they differ. |
| **Assertion already satisfied by the pre-change baseline** | Run the gate with the feature disabled, or at the previous commit. It MUST fail. If it passes, it proves nothing about your change. |
| **Gate reads a stale or generated artifact** | Delete the artifact and re-run. It must regenerate or fail loudly — never pass on the old copy. |
| **Real implementation mocked out** | Assert the production path actually ran: a call count, a side effect, a log line emitted from the real code. |
| **Assertion kind silently ignored** by the harness (unknown key, typo'd field) | Assert something you know is false and confirm the harness reports failure rather than skipping. |
Record the control alongside the pass, in this form:
- `should displace the wave field` — PASS; goes red when `wave_scale` is zeroed
- `web/native WGSL byte-identical` — PASS; goes red when one side is patched by a byte
**A pass with no observed red is reported as UNVERIFIED, not as PASS.**
### Detection methods that actually work
Ranked by observed yield when auditing "green but not integrated" work. CI
suites, PRD checklists, and `done/` placement have caught **none** of it — do
not rely on them.
1. **Grep for a live caller.** For each new symbol, list non-test consumers. One
read-only pass over ten subsystems found ~30 unwired features.
2. **Run the gate's assertion against an unmodified baseline.** If the untouched
starting state passes, the gate measures nothing.
3. **Read the raw log/trace, not the verdict.** The verdict said 3/3 scenarios
pass; the effect log showed the same entity re-emitting `despawn` for 234
ticks and never entering the rendered set.
4. **Drive the real transport/UI, then inspect the resulting state.** A tool
returning `ok, changed: true` had written an empty object.
5. **Look at the output with your own eyes.** Six genre presets produced
indistinguishable arenas; all six automated metrics passed them.
### Phase Verification Template
Each phase MUST include a **Verification Plan**:
```markdown
**Verification Plan:**
1. **Unit Tests:**
- File: `tests/unit/feature.spec.ts`
- Tests: `should X when Y`, `should handle Z error`
2. **Integration Test:**
- File: `tests/integration/feature.int.spec.ts`
- Tests: `should persist data correctly`, `should rollback on failure`
3. **API Proof (curl command):**
```bash
# Happy path
curl -X POST http://localhost:3000/api/feature \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"input": "test"}' | jq .
# Expected: {"success": true, "id": "..."}
# Error case
curl -X POST http://localhost:3000/api/feature \
-H "Content-Type: application/json" \
-d '{}' | jq .
# Expected: {"error": "Unauthorized", "code": 401}
```
4. **Playwright Verification:**
- File: `tests/e2e/feature.spec.ts`
- Flow: Login → Navigate → Action → Assert result
5. **Integration Proof (required, and not satisfied by any test above):**
```bash
# 1. Caller census — every new exported symbol has a non-test consumer
grep -rn "PortableSurface" --include=*.rs --include=*.ts | grep -v "/tests\?/" | grep -v ".spec." | grep -v ".test."
# Expected: at least one hit that is not the definition itself
# 2. Revert check — removing the new path must break something pre-existing
# (rename the symbol / flip the flag off, then run the existing suite)
# Expected: a PRE-EXISTING test or flow fails
# 3. Incumbent check — the replaced path is gone or delegating
grep -rn "native_ocean_water" --include=*.rs
# Expected: no live references, or only a delegation
```
6. **Evidence Required:**
- [ ] All tests pass (`yarn test` / project equivalent)
- [ ] Each gate has an observed negative control (recorded red)
- [ ] curl commands return expected responses
- [ ] E2E test demonstrates full user flow
- [ ] Integration Proof commands produce the expected output (pasted, not summarized)
- [ ] `yarn verify` passes
```
### Verification Checklist by Feature Type
**API Endpoint:**
- [ ] Unit test for request validation
- [ ] Integration test for business logic
- [ ] curl command with expected response documented
- [ ] Error cases tested (400, 401, 403, 404, 500)
- [ ] Rate limiting verified (if applicable)
**Database Change:**
- [ ] Migration runs without error
- [ ] Rollback works
- [ ] Data integrity constraints tested
- [ ] Query performance acceptable (EXPLAIN ANALYZE for complex queries)
**UI Feature:**
- [ ] Component renders correctly (unit/snapshot test)
- [ ] User flow works E2E (Playwright)
- [ ] Loading states handled
- [ ] Error states handled
- [ ] Responsive behavior verified
**Background Job/Cron:**
- [ ] Job executes successfully
- [ ] Failure handling tested
- [ ] Idempotency verified (safe to re-run)
- [ ] Logs show expected output
**Webhook/Integration:**
- [ ] Incoming payload validated
- [ ] Signature verification tested (if applicable)
- [ ] Retry behavior documented
- [ ] curl command to simulate webhook
### Test Naming Convention
`should [expected behavior] when [condition]`
Examples:
- `should return 401 when token is missing`
- `should create user when valid data provided`
- `should rollback transaction when payment fails`
### Evidence Documentation
For MEDIUM/HIGH complexity, include a **Verification Evidence** section in the PRD after implementation:
```markdown
## Verification Evidence
### Phase 1: User Authentication
- Unit tests: 12 passing (screenshot/output)
- curl test: POST /api/auth/login returns JWT ✓
- Playwright: Login flow completes in 2.3s ✓
- yarn verify: PASS
### Phase 2: Dashboard
- Component tests: 8 passing
- E2E: Dashboard loads with user data ✓
- Performance: LCP < 2.5s ✓
```
**Remember: If you can't prove it works, it doesn't work.**
---
## 7. Acceptance Criteria
### Write criteria about the consumer, never about the artifact
This is the single wording choice that decides whether a PRD can pass while
shipping nothing. Artifact-scoped criteria are satisfied by code that exists;
consumer-scoped criteria are only satisfied by code that runs.
| Artifact-scoped (rejected) | Consumer-scoped (required) |
|---|---|
| "the generated shader validates under naga" | "the ocean renders from the generated shader in both runtimes" |
| "7 presets proved" | "each preset produces a playfield distinguishable from the bare starter" |
| "the endpoint returns 200" | "the invoice appears in the user's billing list after checkout" |
| "touch readers are implemented" | "dragging on a touch device moves the player" |
| "the exporter round-trips a shader" | "the shader the product actually uses is exported and consumed" |
| "a preset ships for this genre" | "this genre's reference capture matches within threshold" |
Litmus: could this criterion be checked green by a build that a user could not
tell apart from the previous one? Then rewrite it.
### Never file a PRD as done with unchecked boxes
A PRD moved to `done/` with unresolved boxes makes the whole `done/` directory
untrustworthy as a record. Either the box is checked with evidence, or the PRD
stays open with the gap named.
Binary done checks:
- [ ] All phases complete
- [ ] All specified tests pass
- [ ] `yarn verify` passes
- [ ] All automated checkpoint reviews passed (manual also passed if required)
- [ ] UI exists for user-facing features (or explicitly marked internal-only)
**Integration gates (a PRD with any of these unchecked is NOT done):**
- [ ] Integration Ledger has zero `TBD` cells; every live caller is a real non-test `file:line`
- [ ] Every new exported symbol has at least one non-test consumer (caller census pasted)
- [ ] Revert check passed: disabling the new code breaks a pre-existing test or flow
- [ ] Every `Replaces` row's old path is deleted or delegating — no behavior has two live implementations
- [ ] Every gate has a negative control that was observed failing
- [ ] The capability was proved on the real production subject, or the remaining gaps are listed with their closing phase
---
## Quick Reference
### Vertical Slice (Good) vs Horizontal Layer (Bad)
| Good Phase | Bad Phase |
| -------------------------------- | -------------------- |
| One endpoint returning real data | All types and DTOs |
| One socket event working e2e | All socket handlers |
| One button doing one action | Entire backend layer |
**Litmus test:** Can you describe it as "User does X → sees Y"?
### Anti-Patterns
- Implementing multiple phases without checkpoints
- Phases with no user-testable outcome
- "yarn tsc passes" as sole verification
- Touching 10+ files in one phase
- Skipping automated review when available
- **Backend without UI** - user-facing features with no way for users to access them
### Isolation Anti-Patterns
These are the concrete diff signatures of "implemented but not integrated."
Each one has shipped a fully green PRD that changed nothing for users. Scan the
diff for them at every checkpoint.
| Smell | What it looks like in the diff |
|---|---|
| **Orphan module** | New file whose only importers are its own tests — or zero importers at all |
| **Additive migration** | The new implementation lands and the old one is still the one running. Two or three copies of the behavior, none sharing a source. |
| **Dead-code marker** | `#![allow(dead_code)]`, `eslint-disable no-unused`, unused-export suppression added so the new code compiles |
| **Unread contract** | A types/contract/schema/descriptor file the implementation never consults |
| **Listed-but-absent test** | A test name promised in the PRD with no body in the repo |
| **Uncompiled test** | A test file the build excludes: missing `mod`/import, excluded target, non-matching glob |
| **Self-comparison** | A differential or parity gate whose two sides resolve to the same module, file, or artifact |
| **Toy proof** | The capability proved on the one input that needs none of the hard requirements |
| **Twin constants** | PRD says "derived from one owner"; the code has two hardcoded literals with nothing tying them |
| **Registered but unspawned** | Plugin/handler/route registered, nothing ever invokes it |
| **Manufactured evidence** | The report emits `status: "applied"` / `ok: true` as a literal instead of measuring anything |
| **Vacuous fixture** | The gate's fixture does not contain the feature under test (an overlay-packaging gate whose fixture has no overlays) |
| **Envelope ≠ state** | The call returns success and the persisted state is unchanged — `changed: true` written next to an empty object |
| **Pure function stands in for the loop** | The evidence harness calls the function directly; the frame loop / request path never does |
**Rule:** finding any of these at a checkpoint fails the phase. Fix the wiring
in the same phase — never log it as follow-up.
## Principles
- **SRP, KISS, DRY, YAGNI** - Always
- **Composition > inheritance**
- **Explicit errors** - No silent failures
- **Automated verification** - Let the agent catch drift
---
## Checkpoint handoff
After each phase, record the complete evidence packet named by the Checkpoint
Protocol: the exact commands, the observed-red result for every gate, caller
census, revert check, and any remaining blocker. Hand the packet to the manager
and stop. The manager chooses the single read-only review path described by
`references/runtime.md`; this skill never starts that process and never chains
execution automatically.
prd-swarm-coordinator23.8 KB
--- name: prd-swarm-coordinator description: Execute one or many conforming PRDs through contract-preserving Codex lanes with bounded parallelism, inherited gates, review, and honest delivery states. license: MIT --- # PRD Swarm Coordinator ## Scope and runtime contract Own one or many PRDs through the same intake, brief, scheduling, worker, review, repair, and delivery path. A single PRD is a swarm of one; it does not take a shortcut. The `references/` directory is at the plugin root, beside `skills/`; from this file resolve it as `../../references/`. Read `references/prd-contract.md` and `references/intake.md` before intake, and read `references/runtime.md` before any preflight or delegation. The executable checks live in `scripts/linchpin.sh`. The manager is the current session's Manager role. Workers, repair workers, integration workers, and conflict workers use the Worker role only through the `codex exec` subprocess shape in `references/runtime.md`. The reviewer is a fresh `codex exec --sandbox read-only` process at the Reviewer row's model and effort. Read those values from the table; an effort written into this sentence is a second copy that goes stale the first time the pin changes. Never use a native subagent for Luna, never change tier after a failed attempt, and use `codex exec resume <session-id>` only for a recorded continuation. This skill does not support generic non-PRD swarm requests, and it does not arm the optional goal loop. Use the referenced documents and `scripts/linchpin.sh` subcommands as interfaces: invoke the specific check you need and inspect its output; do not read the full helper source into context. ## Intake branch 1. Read the complete input artifact. Do not summarize it, and do not rewrite it. 2. **Execute the PRD the user pointed at, as written.** A missing `prd_contract: v1` marker, a legacy heading, a prose file list, or an absent ledger does not block execution and is not a reason to migrate, re-author, or draft a replacement. `scripts/linchpin.sh brief <prd>` transfers whatever sections exist verbatim and marks the rest `NOT DECLARED`; the worker follows the PRD's own phases and file lists from there. Standardize only when the user asks. The only blocker is a path that is not on disk. 3. Preserve whatever the artifact does declare — Integration Ledger, Negative Controls, Acceptance Criteria, Checkpoint Protocol — verbatim. Do not re-derive a shorter checklist, and do not add a gate the author never asked for to compensate for a section the PRD does not have. 4. A creator output never auto-starts this skill. Require explicit confirmation after creation or upgrade and before preflight. 5. Use `references/intake.md` for intent, complexity-floor refusal, config, and capability routing. Repository state cannot override a direct write-PRD request. ## Preflight and configuration Read optional `.linchpin.toml` using the defaults and validation in `references/intake.md`. Its absence is valid. Record the resolved values in the run ledger and persist typed natural-language overrides before scheduling. Run these checks before branches or workers: - verify the target is a Git repository; - run `scripts/linchpin.sh preflight` against `$CODEX_HOME/models_cache.json`; it resolves both role models and confirms `$CODEX_HOME` is writable, which is what every worker and reviewer subprocess needs before its model starts; - inspect current status, default branch, remotes, and delivery capability; - parse every phase `Files (N)` list; a malformed list is an error, never an assumption of disjointness; - verify the review setting was explicitly chosen if it is false. Only a missing Git repository or missing worker capability is a refusal. A missing worktree, dirty unstashable tree, missing remote, or missing PR client is a named degradation unless the user explicitly forced the unavailable mode. ## Contract-preserving worker brief Generate each brief with the resolved lane metadata, writing it to a file: `scripts/linchpin.sh brief <prd> <lane-id> <lane-mode> <delivery-mode> --config-dir <target-repo> --out <brief-file>`, then verify it with `scripts/linchpin.sh brief-check <prd> <brief-file> --config-dir <target-repo>` and pass that file's contents as the worker prompt. Pass `--config-dir` to both: your working directory is not necessarily the target repository, and a brief emitted with the repository's config but checked without it fails its own verification on a stale runtime pin. The brief is the handoff; a prompt you compose yourself instead is a dropped ledger and a dropped scope rule. Resolve these values before invocation from `.linchpin.toml`, the file-intersection group, and delivery capability. With no config file, the helper's `brief <prd>` form remains a lane-1/parallel/pr default for direct callers; production lanes pass all three resolved values. The brief contains, in this order: 1. source PRD path and lane identity; 2. every parsed file path from every phase; 3. the complete Integration Ledger copied verbatim, including every row's Live caller and Negative control; 4. the complete Negative Controls table copied verbatim; 5. the complete Acceptance Criteria and Checkpoint Protocol copied verbatim; 6. the runtime-derived worker/reviewer invocation shapes, lane mode, delivery mode, and prohibited actions. Model and mechanism come only from `references/runtime.md`. Effort comes from there too unless the target repository's `.linchpin.toml` sets `worker_effort` or `reviewer_effort`, which is why the brief is generated with `--config-dir`. Read the emitted values out of the brief rather than restating a pin from memory. Before launch, compare ledger row ids between source and brief. A missing row, caller, or control rejects the brief. The worker must not be asked to infer missing acceptance criteria from a summary. ## Per-group mode selection Run `scripts/linchpin.sh mode <resolved-execution> [--config-dir <target-repo>] <prd...>` after all lists parse. It builds the file-intersection graph and emits one group per connected component: - disjoint groups use parallel worktrees when worktree creation succeeds; - groups with intersecting file sets run sequentially, one lane at a time; - a PRD with no `Files (N)` list has its set derived from its prose `**Files:**` paragraphs, for grouping only — the file on disk is never rewritten; - a PRD that declares no file set at all takes its own group with its isolation announced as unproven; it never drags the rest of the batch into its queue; - explicit sequential mode makes all groups sequential; - explicit parallel mode fails loudly on intersection or worktree failure; - auto mode degrades only the affected group and announces the reason; - `max_lanes` is a real concurrency bound; each group reports `active=` and `queued=` lanes when capacity is exceeded; - one lane uses the same output, gate, review, and delivery fields as any other group and has no special branch. When a group must degrade, run `scripts/linchpin.sh schedule auto <status> [--config-dir <target-repo>] ...` with the status that actually happened — `worktree-fail`, `dirty-tree`, `unparsed-files`, or `config`. Attempt the real `git worktree add` before you claim it failed; the announcement the user reads must name the true reason. Announce the sequential fallback before starting the first lane, and preserve the same brief and gates. The schedule output identifies active and queued lanes under `max_lanes`. Never abort a normal auto run for unavailable isolation. For a forced parallel run, fail with the exact capability error so the user can correct the environment. Mode is per group. A colliding pair may be sequential while an independent pair remains parallel. A worker never receives a weaker gate because its group is sequential. ## Lane lifecycle Run `scripts/linchpin.sh workspace <target-repo>` **before the first write to `.linchpin/`**, including the ledger. It creates the directory and adds `.linchpin/` and `.worktrees/` to that repository's `.git/info/exclude` unless they are already ignored. Run output is Linchpin's scratch space, not the user's work: it must never appear in `git status`, never be staged into a lane commit, and never force a manual cleanup after the batch. The entries go in `.git/info/exclude` rather than `.gitignore` on purpose — ignoring our own output must not itself leave a modified tracked file behind. If the user asks for the ignore to be committed instead, add `.linchpin/` to `.gitignore` and say so; do not do it unasked. Keep a run ledger at `.linchpin/run-<timestamp>.md` in the target repository, written before the first worker starts and updated as each lane changes state. Write every row with the helper, never by hand: ```sh scripts/linchpin.sh lane .linchpin/run-<timestamp>.md <lane-id> \ --set state=RUNNING --set prd=<path> --set branch=<branch> --set pid=<pid> ``` Each call upserts that lane's row and keeps the fields earlier calls wrote, so record what you know when you know it. For every lane record the PRD, slug, baseline, branch, worktree or shared-tree mode, file set, overlap group, dependencies, process id, subprocess session id, brief path, verification commands, review state, repair rounds, delivery mode, and terminal evidence. `lane` refuses a row it cannot verify: an unknown state, `MERGED` as a product state, a `commit` sha that does not resolve in the repository, a `DELIVERED(...)` row missing its prd, branch, commit, gates, or review, a `gates` path that is not on disk, or a `BLOCKED` row with no `reason` and `resume`. That refusal is the point — a lane recorded as committed whose sha the worker never created is the false ledger row this run exists to make impossible, and it is not a claim you can talk your way past. Fix the row or fix the lane. A run with no ledger file on disk is not resumable, and an unresumable run is not a run. 1. **Every lane gets its own branch**, sequential ones included: `git switch -c linchpin/<lane-slug>` from the same detected base branch. Never branch a lane from another lane, and never let a worker commit onto the branch the user had checked out. Sequential means one lane at a time in the shared tree; it never means committing onto the user's working branch. Branch from the **remote** base after a fetch (`origin/<base>`), not from the local branch of the same name. A local base that sits ahead of its remote carries the user's unrelated committed work into every lane, and delivery then merges that work under a PR title that never mentions it. If local and remote have diverged, say so before the first lane starts; whose commits those are is the user's call, not a detail to resolve silently. For a worktree lane, create it with `scripts/linchpin.sh worktree <main-repo> <lane-slug> <base>`, always from the repository's **main** worktree and never from inside another lane. It fetches, resolves `origin/<base>`, refuses a nested or already-claimed lane, and prints `WORKTREE-READY` or a `WORKTREE-FAIL <reason>` you pass straight to `schedule auto`. Do not substitute a worktree helper from the user's machine: one of those pulled the base branch inside the user's dirty source tree on the way to creating the worktree and left an unresolved merge across twenty uncommitted files. Lane isolation is Linchpin's to perform, and the source tree the user is sitting in is never modified to create a lane. 2. **Make the worktree able to run the gates before the worker starts.** A fresh worktree has source but no build state: no installed dependencies, and none of the repository's local tooling. Every lane that discovers this alone discovers it again in parallel, and reports a gate it could not run as if that were a verification result. Once per run, resolve the bootstrap for this repository — its lockfile install, its pinned runtime version, and whether dependencies can be shared across lanes rather than installed per lane — then apply it to each worktree and state in the ledger which gate commands are actually runnable there. Check the repository's own test configuration for path exclusions that would silently match your worktree directory and match zero tests; if one does, resolve the override once and put it in every brief rather than letting each lane rediscover it. The same gap breaks commits: a repository whose commit hooks run out of `node_modules` rejects every commit in a fresh worktree until dependencies are installed, so bootstrap before the first commit rather than after it. If what you are committing is your own manager housekeeping — a ledger, a plan file — recording it with hooks skipped is acceptable and worth saying out loud; a lane commit is never delivered on skipped hooks. 3. Launch workers using only the Worker row in `references/runtime.md`. Pin the required effort and working directory in the subprocess invocation, and pass the generated brief file as the prompt. Do not inherit session defaults, do not retype the brief into a prompt of your own, and do not route code edits through another runtime. 4. Require a worker commit, exact test output, caller census, revert check, and gate evidence before manager verification. The commit is evidenced by the commit itself, never by a worker's summary claiming one. A lane recorded as committed whose sha the worker never created is a false ledger row. 5. Keep a partial lane and its worktree. A timeout or worker summary is not a delivery result; inspect the actual diff and resume from the recorded state. 6. A lane that ends `PARTIAL` or `BLOCKED` releases its group's queue. The next queued lane starts; the batch does not stall behind a lane that is done failing. ### Awaiting a lane A lane takes minutes, and its progress prose is not evidence you will act on. Launch each worker **detached**, with its output redirected to a log and its process id written to `.linchpin/<lane>.pid`, so that no lane holds an interactive session open for you to babysit: ```sh codex exec ... > .linchpin/<lane>.log 2>&1 & echo $! > .linchpin/<lane>.pid ``` Then wait on the whole group at once: ```sh scripts/linchpin.sh await .linchpin/<lane-a>.pid .linchpin/<lane-b>.pid --interval 60 ``` It blocks until every lane in the group has exited and prints one `AWAIT-DONE` row per lane. Waiting for a group in one call costs turns in proportion to the number of groups; polling each live subprocess on a short timer costs turns in proportion to lane duration, and runs that did it spent hundreds of turns restating that a lane was still running. Process exit and the real diff are the only two signals worth a turn; announce a lane's status when it changes, not on a timer. ### Ending the run The run is over when the workspace looks the way it did before it started. For every lane, after its delivery state is terminal and recorded: remove the worktree, prune the worktree list, and delete the lane branch that was merged. Keep the worktree and branch of any lane that ended `PARTIAL` or `BLOCKED` — those are resumable state — and name in the final report exactly what was kept and the command that resumes or removes it. Leave `.linchpin/` in place as the run record; `workspace` has already kept it out of `git status`. Check the target repository's status at the end and account for anything Linchpin added that is still there. Cleanup is part of delivery, not an optional courtesy. ## Inherited lane gates The PRD's Negative Controls table is an inherited lane gate, not advice. Copy it into the reviewer packet with the ledger and require one Gate Evidence row for every control. The evidence format is: ```markdown ## Gate Evidence | Gate | Result | Observed-red evidence | Exact command/result | |---|---|---|---| | gate-id | PASS | RED observed: disabled gate | `command: sh tests/example.sh`; result: RED observed: disabled gate; exit: 1 | ``` The exact command/result cell repeats the command documented in the PRD's Negative Controls table. `gate` rejects missing, duplicate, or extra gate ids, generic evidence without that exact command, green-only evidence, and zero exits. Run `scripts/linchpin.sh gate <prd> <report>` before delivery. A report with only green assertions, a missing control, or a missing observed-red line is `UNVERIFIED` and rejected. Every control must have failed as expected at least once. This rule is identical in parallel and sequential mode. When the PRD declares no Negative Controls, `gate` reports `GATES-NOT-DECLARED` and delivery proceeds on the verification the PRD *does* declare. Do not invent controls the author never wrote, and do not hold a lane because a section is absent. The inherited-gate rule binds the controls a PRD declares; it never manufactures new ones. The reviewer packet must contain the negative-control table even when all functional tests are green. The manager records the exact red command and its non-zero result; a verbal claim is not evidence. ## One review and repair rule Two preconditions come before the reviewer, in this order. A lane that fails either is not ready for review, and launching a reviewer anyway produces a rejection that says nothing about the code: 1. **The lane is committed.** An uncommitted working tree is `PARTIAL`. Get the worker's own commit first; a missing commit is a worker-contract failure that no review round can fix. 2. **You have run the gates yourself.** The reviewer is `--sandbox read-only`: it cannot install dependencies, write a cache, bind a port, or run the repository's suites. Run them in a writable tree and produce the Gate Evidence table before the reviewer starts. Generate the review brief with the helper; it refuses to emit without both: ```sh scripts/linchpin.sh review-brief <prd> <lane-id> --gates <gate-evidence.md> --commit <sha> --out <review> ``` Then launch exactly one fresh reviewer per lane through this shape, with all role values resolved from the Reviewer row in `references/runtime.md`: ```text codex exec --model <Reviewer.Model> -c 'model_reasoning_effort="<Reviewer.Effort>"' --sandbox read-only -C <lane> "$(cat <review-file>)" ``` Pass the file `review-brief --out` wrote. Do not interpolate the packet's text into the command line: it contains backticks, quoted commands, and table pipes, and a manager that hand-escaped one sent `codex` a mangled argument and read its `Reading additional input from stdin...` as a review. If the reviewer exits before the model starts, that is an environment failure, not a verdict — the usual cause is a `$CODEX_HOME` the process cannot write, which preflight already checks. Report it as an unresolved external gate and say the lane is unreviewed; never let a reviewer that could not start be recorded as a lane that passed. Record `review_used: true` before launch. The reviewer cannot edit. The manager closes findings after Luna repair; no second reviewer is started after repair. Every finding is labelled `DEFECT` or `EVIDENCE-GAP`. Only a `DEFECT` blocks delivery. An `EVIDENCE-GAP` is recorded in the ledger and delivered past — a reviewer reporting what it was structurally unable to run is describing its own sandbox, not a fault in the lane. State facts in the brief without classifying them. A brief that says "treat the missing commit as a finding" has already written the verdict, and the review that comes back is an echo. `APPROVE` with zero findings is a valid review. What this review is for is the class of defect only a reader reaches: a negative control that stays green when the feature is deleted, a field the code accepts and never maps, a document asserting behavior the code contradicts. Narrow the question to that; never trade away the rigor. When a worker or reviewer exposes a failure, treat it first as specification evidence. Before re-delegating, write a corrected or narrowed handoff naming the exact file, line, expected behavior, failing command, and newly required test. The handoff must differ from the failed prompt. Never repeat an unchanged prompt, and never change the model tier or effort to escape a failed gate. Use the recorded `codex exec resume <session-id>` only when the corrected continuation is explicit. A repair round is for a `DEFECT`. Spawning a fresh worker whose entire scope is "commit the diff that already exists" is not repair — it is manager integration work, and it costs a full model run to reach a `git commit` you could have made directly. If a lane arrives uncommitted, fix it as integration and record the worker-contract failure; do not dress it up as a repair round. A lane whose PRD requires that nothing be committed is already correct and never gets one. Repair, integration, and conflict work face the same inherited gates. Ordinary overlap and merge conflicts are manager-directed integration work; preserve the accepted intent of both PRDs and add a regression test when resolution combines behavior. A semantic conflict that would require violating a PRD is a named `BLOCKED` decision, not a silent rewrite. ## Delivery and terminal states Resolve delivery from `.linchpin.toml`: `pr` by default, `branch` when selected, and `branch` after an announced missing-remote or missing-client fallback. A delivery fallback does not remove review, gate evidence, or verification. Probe the PR path once, at preflight, rather than discovering it lane by lane at the moment of delivery. Confirm which client actually works against this remote, and which merge methods the repository permits, before the first merge — a repository that forbids merge commits rejects the merge after the PR is already open. Record the working client and the permitted method in the ledger and use them for every lane. **Merging to a shared base branch is the one stop-and-confirm point.** Opening PRs, pushing lane branches, and reporting are all Linchpin's to do. Merging another lane's work into a branch other people build on is the user's decision, and an autonomous run is exactly the situation where nobody is watching. Ask once, before the first merge, and carry the answer across the remaining lanes. If the run was told not to stop, deliver every lane as an open PR and say that the merges are waiting — an unmerged PR is recoverable, a merge is not. Use only these terminal forms: - `DELIVERED(pr)` or `DELIVERED(branch)` after the full evidence packet passes; - `BLOCKED <named external reason> <resumable command>` with preserved state; - `PARTIAL` while implementation or evidence is incomplete. Never label a lane `MERGED` as its product state. Never call a lane delivered because a process exited, a summary said done, or a green-only test suite ran. ## Final controller verification Before handing off, inspect the real branch or shared-tree diff and run the gates the **target repository** names — its own test, lint, typecheck, and build commands, discovered from that repository. Linchpin's own repository scripts (`scripts/verify.sh`, its shellcheck and jq checks) are for developing this plugin; never run them inside a user's repository. Confirm the diff contains only the files the PRD's scope covers. An unrelated deletion, an unrelated dependency bump, or an unrelated doc edit that arrived inside a lane commit is a finding, not a bonus: name it, and get the worker's own commit narrowed before delivery. Read the ledger back rather than recalling it: ```sh scripts/linchpin.sh status .linchpin/run-<timestamp>.md ``` It prints one line per lane and exits `0` only when every lane is `DELIVERED(...)`, `1` while any lane is still open, and `2` when the only unfinished lanes are `BLOCKED`. Summarizing eight lanes from memory at the end of a long batch is where a run starts reporting work it did not do; the command is the answer, and its exit code is the honest one. The final report maps every PRD criterion and every ledger row to a command, file:line, or captured result. It names observed-red failures, unresolved external gates, and a resumable command. It never claims the external install swap or live post-swap resolution without owner evidence.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- MIT
- Package author
- Linchpin contributors
- Keywords
- See publisher keywords
Declared capabilities
- Interactive
- Write
Package observed Oct 3, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 3, 2026 · 06:00 UTC
- Collection status
- Collected
plugins_6a72db7ab4408191a9da4d85532ebe54
Download plugin data (JSON)