← Files okrdevARCHIVED FILE

docs/stack.md

20.7 KB · Oct 2, 2026 · 00:30 UTC

↓ Download file

# The stack module

One good path from idea to production, built so anyone on the team — technical or not, human or
AI — can ship safely. Next.js on Vercel. Neon Postgres with a database branch per preview. Drizzle
migrations. Neon Auth. Vitest and Playwright. GitHub Actions. A protected `main` where merging a
PR *is* the deploy. Feature flags as code in the repo.

**You do not need any of this to run okrdev.** The method — parking lot, cycles, check-ins,
retros, the coach — installs on any stack, or on no stack at all. Existing codebase? Keep it;
[adoption.md](adoption.md) covers brownfield installs. The stack module is for greenfield
projects, or for teams making a deliberate migration with eyes open. `/okrdev:install` offers it
only in those two cases, and never as a prerequisite for anything above it on the ladder.

## The loop it buys you

Every change follows the same path, and no change follows any other path:

1. A builder (a human working with AI) opens a pull request.
2. Vercel deploys a **preview** — a private copy of the app at its own URL — and Neon forks a
   **database branch** for it from a parent seeded with synthetic data. Never from production.
3. CI lints, typechecks, and runs unit tests. The PR's committed migrations are applied to the
   preview's database branch, then Playwright smoke tests run against the preview URL with a
   seeded test user.
4. A human clicks the preview link and checks that the change does what it claims. No code
   reading required — [shipping-explained.md](shipping-explained.md) teaches the vocabulary.
5. Review lands. Where the paths are risky (migrations, auth, payments), CODEOWNERS pulls in the
   domain expert automatically. Squash merge.
6. Merge deploys production. A gated Actions job applies the same migration files to the
   production database. The preview and its database branch are deleted.

There is no deploy button, no staging environment to babysit, no "works on my machine." The PR is
the unit of change, the preview is the proof, and `main` is the truth.

## Why one opinionated path

The stack is chosen to be **AI-legible** — the property that matters most when agents do most of
the building:

- **The whole system is text in the repo.** Schema, migrations, tests, CI, deployment behavior —
  all files an agent can read, grep, and modify through the same PR flow as everything else.
  Nothing important lives only in a dashboard.
- **No mid-flight decisions.** One database, one ORM, one test runner, one way to deploy. An
  agent (or a new teammate) never has to ask "which of our three ways do we do this here?"
- **Boring, popular tools on purpose.** Next.js, Postgres, Playwright — these have years of
  documentation and deep training-data coverage. Agents hallucinate less on roads well traveled.
- **Ground truth at every layer.** A preview URL to click, a CI verdict to read, a database
  branch to query. The agent and the non-technical DRI verify with the same tools, which is what
  makes "anyone can own anything" ([roles.md](roles.md)) more than a slogan.
- **Safety is structural.** Untested code can't merge, unreviewed risky paths can't merge,
  production can't be force-pushed. Nobody has to watch. Rails, not vigilance.

## The choices

Each choice states what it is, why it won, and how you would leave. A stack you can't leave is a
stack you can't trust — every exit path below is documented on purpose.

### Hosting: Vercel

**What.** Next.js (App Router, TypeScript) deployed on Vercel. Every PR gets a preview
deployment at its own URL; every merge to `main` deploys production.

**Why.** Previews are the heart of the stack: they give non-technical DRIs a real, clickable
copy of the app for every proposed change, with zero setup. Vercel makes that the default
behavior rather than an infrastructure project.

**Exit.** The app is standard Next.js — `next build` runs on any Node host or in a container.
What you'd rebuild elsewhere is the per-PR preview pipeline, and that's the piece you'll miss
most. Avoid Vercel-only primitives where a portable option exists and the exit stays cheap.

### Database: Neon Postgres, one branch per preview

**What.** Neon Postgres, wired to Vercel through the Neon integration. Each preview deployment
gets its own copy-on-write database branch, forked from a **seeded parent branch containing
synthetic data — never production**. Branches are deleted when the PR closes; a scheduled
workflow ([../templates/github/workflows/neon-cleanup.yml](../templates/github/workflows/neon-cleanup.yml))
sweeps up orphans.

**Why.** A preview with a shared database is a lie — click-testing against data another PR just
mutated proves nothing. Branch-per-preview makes every preview a genuinely isolated app-plus-data
environment. And the seeded-parent rule is a hard privacy line: preview links get pasted into
chat, opened on phones, and shared with people who have no production access. A preview forked
from production is a data leak with a URL. PII never enters a preview, so a leaked preview link
leaks nothing.

Two operational notes, because they bite at the worst time: verify delete-on-close in the
integration settings (don't assume it), and know your branch limit — Neon's free tier allows
around ten branches per project, which is one busy week of PRs. Details sit next to the setup
steps in [../templates/stack/README.md](../templates/stack/README.md).

**Exit.** It's Postgres. `pg_dump`, restore anywhere, done. What you lose is instant
copy-on-write branching — which is the reason to be here in the first place.

### Schema and migrations: Drizzle

**What.** Drizzle ORM with the schema defined in TypeScript and migrations generated as plain
SQL files, committed to the repo.

**Why.** The schema-as-TypeScript is agent-readable and agent-writable. The generated SQL is
human-reviewable: a schema change shows up in the PR diff as the exact DDL that will run, which
is what lets CODEOWNERS put a domain expert in front of every migration.

The runbook, in one table:

| When | Command | Against |
|------|---------|---------|
| Local experimentation | `drizzle-kit push` | your own dev branch only |
| Every schema change | `drizzle-kit generate`, commit the SQL | the repo, in the PR |
| CI, per PR | `drizzle-kit migrate` | the PR's preview database branch, before Playwright |
| Production | `drizzle-kit migrate` in a gated Actions job | production, on merge |

`push` never touches a shared branch — it mutates the database directly with no committed
artifact, which means no review, no replay, no audit trail. Everything shared goes through
generated, committed SQL.

Breaking changes use **expand/contract**, because deploy = merge means old code and new schema
briefly coexist: (1) expand — add the new column or table alongside the old, backward-compatible;
(2) dual-write and backfill; (3) switch reads to the new shape; (4) contract — drop the old in a
later PR once nothing reads it. Each step is its own PR, so every commit on `main` runs against
the schema it finds.

**Exit.** The migrations are plain SQL files — any tool, including `psql`, can replay them. The
only Drizzle-specific artifact is the TypeScript schema in your app code.

### Auth: Neon Auth

**What.** Neon's managed auth, with one property that made the decision: user records sync into
a `neon_auth` schema **in your own Postgres database**. Your users are rows you can join
against, back up, and export.

**Why.** Auth vendors that keep your users in *their* database hold your business hostage at
exit. Here, `pg_dump` gets your user table any day of the week. Managed auth also removes the
single riskiest thing a generic builder could hand-roll — password handling.

**Exit.** User records: already yours, export at will. Password hashes live with the provider,
so leaving means exporting hashes where supported or running a password-reset campaign — plan
for the second and be pleasantly surprised by the first. If a hosted auth dependency is
unacceptable from day one, the sanctioned fallback is Auth.js, self-hosted against the same
Postgres. Nothing else in the stack changes.

### Tests: Vitest and Playwright

**What.** Vitest for unit tests, run in CI on every PR. Playwright for smoke tests, run **against
the real preview deployment** — same build, same database branch, same auth as what would ship —
signed in as a seeded test user, passing the `x-vercel-protection-bypass` header so deployment
protection stays on for everyone else.

**Why.** Unit tests catch logic mistakes cheaply. Smoke tests against the preview catch the
mistakes that matter — the thing you verify is the thing you ship, not a localhost approximation
with mocked data. Keep the smoke suite small and meaningful: it runs on every PR against real
infrastructure, so every test has to earn its seconds.

**How tests get written — two rules, both halves of okrdev's own method:**

- **Red first, where it earns its cost.** For a substantive fix or a load-bearing path, the
  bug becomes a failing test before the fix, and the fix is done when it turns green. The
  coach does the building, so the coach writes the failing test silently when one is cheap —
  "I made the check fail first, on purpose, so we know the fix worked" — and simply skips it
  when the repo cannot honestly assert the behavior. No ceremony lands on small fixes:
  maintenance still classifies silently, and a coach that makes bugfixes feel expensive
  teaches people to stop mentioning bugfixes. That never-list line outranks this rule
  everywhere they touch.
- **Promotion is chafe-gated, not automatic.** The How-to-verify steps on a PR are already a
  test a human runs — [evidence.md](evidence.md) calls them the domain language made
  falsifiable. They get promoted into a Playwright smoke test against the preview **when the
  path has broken once or burned a DRI** — not one test per shipped capability. A smoke
  suite grown from real breakage stays small, meaningful, and trusted; a smoke suite grown
  from completeness becomes the untended garden nobody believes when it goes red.

**The loop is local; CI confirms rather than discovers.** Red-first is priced by the loop
that runs it: an easy sell when lint, typecheck, and the Vitest units answer in seconds on a
laptop, and a hard one when every answer costs a push and a wait. So run the cheap
deterministic checks where they are cheap — the preview still owns verification, branch
protection still owns enforcement, and anything you wire locally is opt-in and step-overable,
because a knowingly-red push has to stay first-class. One warning travels with the speed: a
local green that skipped a missing linter, or ran on a different runtime major than CI pins,
is not the green it looks like. okrdev tells itself this with a parity advisory in its own
harness (`check_local_loop` — [testing.md](testing.md) § What runs where); you are told the
failure mode, not handed a script. Nothing here lands in your repo.

A green suite proves less than it feels like it proves: the verdict order stays demo above
suite, metric above demo, exactly as [evidence.md](evidence.md) ranks them. Tests are the
regression floor under real-use evidence, never a substitute for it.

**Exit.** Both are standard open-source tools. There is nothing to exit from.

### Delivery: GitHub Actions, protected main, squash merges, deploy = merge

**What.** CI on GitHub Actions ([../templates/github/workflows/ci.yml](../templates/github/workflows/ci.yml)).
`main` is protected: required CI check, required review with Code Owners, no force pushes,
squash merges only. Merging is the only way anything reaches production.

**Why each piece:**

- **Actions**: CI lives where the code lives, and workflows are files in the repo — AI-legible,
  reviewable, versioned like everything else.
- **Protected main**: the gates are only real if they can't be walked around. Without branch
  protection, CODEOWNERS is a suggestion and required checks are decoration.
- **Squash-only**: one PR becomes exactly one commit on `main`. The PR's `KR:` line rides in the
  squash commit message, so the coach's drift check ([ai-coach.md](ai-coach.md)) can read
  alignment straight out of git history — and reverting a bad change is one clean commit.
- **Deploy = merge**: a separate deploy button is a second path to production, and second paths
  accumulate snowflake state. When merge is the only deploy, git history *is* deployment
  history.

okrdev's own state writes fit inside the gates rather than around them: captures are
`okrdev:parked` issues (zero commits), and the remaining ledger writes — triage results,
check-in files — are batched by ritual and land as small, immediately merged state PRs, about
one a week. The setup script
([../templates/stack/branch-protection.sh](../templates/stack/branch-protection.sh)) still
ships a direct-push bypass for `okrdev/**` writes as an opt-in convenience, commented out by
default — and is honest about its mechanics: GitHub can't scope a bypass to file paths, so
it's actor-scoped, with the `okrdev/**`-only discipline enforced by the coach's contract.

The Level 2 rails — PR template, okr-gate, CODEOWNERS — are part of the method, not the stack,
and remain opt-in ([adoption.md](adoption.md)). The stack simply arrives with everything they
need already in place. Note the plan requirements: branch rulesets are free on public repos;
private repos need GitHub Pro (personal) or Team (organization).

### Feature flags: the Flags SDK, one opt-in brake

**What.** Every feature flag is a typed function declared in `flags.ts` with Vercel's
open-source, provider-agnostic Flags SDK, its default value committed beside it. Flags are
created, changed, and deleted only by PR — the same path as every other change. Percentage
rollouts are committed values too: ramping 5% → 50% is a diff, and deploy = merge makes that
diff live in minutes — provided the percentage hashes a stable identity (the signed-in user
id, or a cookie for anonymous traffic), or the same visitor flips variants on every request.
No flag provider, no dashboard, nothing in the day-one setup — the first flag is one package
and one file. Two rules ride along, and they are yours to wire — a comment convention and a
grep, not a shipped template: every declaration carries a removal criterion a check can
actually read (a date, or "at 100% rollout"), and a CI step you write warns on any flag past
it. Agents mint flags cheaply, and a stale flag is a permanent conditional branch wearing a
temporary one's name.

**Why.** A flag decides what production does — the same authority a migration carries — so
it rides the code's rails or it becomes a second path to production. A vendor dashboard's
flag flip is exactly the mutation the migration runbook bans `drizzle-kit push` for: no
committed artifact, no review, no replay, no audit trail. In `flags.ts` the flip is a diff —
CI runs, review lands, squash merges, and one clean revert is the kill switch, the same one
the stack trusts for every other defect. Previews inherit isolation for free: the flag state
a preview exhibits is the flag state its PR proposes, because it is versioned with the
branch — never a shared vendor environment where another PR's toggle silently changes what
the DRI is click-testing. And the whole surface — every flag, default, call site, and
removal criterion — is text an agent can grep, because there is no dashboard for anything to
live in instead. The brake below adds the one out-of-band store, kill-only, and its whole
discipline exists to keep this paragraph true.

**The brake — opt-in, dormant, kill-only.** The routine kill is a revert: minutes, fully
gated. Where minutes are too many — payments, auth, anything where a bad variant does damage
while CI runs — the one sanctioned escape hatch is a Vercel Global Config (formerly Edge
Config) kill list read at the top of the decide path. Kill-only is the load-bearing
property: an entry can only force a flag to its committed safe value, retreating to behavior
that already shipped through review — it cannot invent behavior that never did. Writes
propagate globally within about ten seconds (Vercel's documented bound), no redeploy. This
is out-of-band state, named as such — the one bounded exception to the loop's "no change
follows any other path," and everything that follows exists to close it. The same shape as
[branch-protection.sh](../templates/stack/branch-protection.sh)'s opt-in bypass, but held to
a stricter discipline: every write is an emergency, logged the moment it happens as a
judgment-call line in the week's check-in file ([ai-coach.md](ai-coach.md)) — or, running
the stack without the method, a dated line in whatever log you keep; a reconciling PR
(change the default, or delete the flag) follows same-day; and a scheduled check you wire
when you arm the brake — the same shape as
[neon-cleanup.yml](../templates/github/workflows/neon-cleanup.yml)'s orphan sweep, because a
PR-triggered check is silent in exactly the quiet week that forgets — fails while the kill
list has been non-empty for more than a day, so "temporarily killed" cannot become a silent,
permanent fork between repo truth and production truth. Install it with your first risky
path, use it never, log it once when you finally do. One operational note: the flip is
instant on dynamically rendered paths and middleware, but statically cached pages hold the
old value until revalidated — put revalidation in the kill runbook, or keep killable flags
off static pages.

**Split tests — deliberately not wired.** This row prescribes no experimentation engine, on
purpose. Below a traffic floor — on the order of hundreds of the KR's *own* conversion events
per week — the power arithmetic refuses everyone equally: detecting anything subtler than a
near-doubling takes more weeks than the cycle has, whichever engine draws the bar, and a
probability bar over a hundred users is [evidence.md](evidence.md)'s rung-1 warning
(measurement dishonesty) rendered as a chart that recruits belief. Below the floor, the
honest instrument is the one the stack already has: ship to everyone, and watch the KR's own
number move week over week through a committed SQL query against your own Postgres. When a KR
clears the floor and a real decision hangs on a test, buy the verdict arithmetic through a
Flags SDK provider adapter — call sites untouched; PostHog is the boring default — because
the statistics that decide whether a KR moved sit beside password handling on this stack's
never-hand-roll list.

**Exit.** Downward there is nothing to exit: flags are code you delete. Sideways, if the SDK
stagnates or folds into a metered product, the prescription survives as the pattern — a typed
flag function over committed defaults is an afternoon to re-implement with zero call-site
changes, and the brake is one KV lookup behind your own code, swappable for any store (or
deleted) in a `flags.ts`-only change. Upward, the same adapter seam is the on-ramp: any
provider, adopted without touching a call site — priced honestly: an adapter hands that
flag's state back to a dashboard, so scope it to the experiment's flags and keep every other
default committed, or this row's argument leaves with it.

## Previews are for humans

The preview link is the one verification tool a non-technical DRI has. If it hits a login wall,
the DRI either rubber-stamps changes or stops merging — both defeat the stack. So preview
accessibility is a **required, verified install step**, not a nice-to-have: deployment
protection stays on, shareable links (or team access) let DRIs open previews without a Vercel
account, and the setup isn't done until someone has opened a preview from a phone or a private
browser window and seen the app, not a login page. The step-by-step is in
[../templates/stack/README.md](../templates/stack/README.md); the review routine a DRI runs
against a preview is in [roles.md](roles.md) and [dri-onboarding.md](dri-onboarding.md).

## Setting it up

Budget a day. The full walkthrough — create-next-app through a verified end-to-end shipping loop
— is [../templates/stack/README.md](../templates/stack/README.md). It installs:

- the app, Vercel project, and Neon integration with branch-per-preview and the seeded parent
- Drizzle and the migration runbook wiring
- Neon Auth and the seeded test user
- Vitest, Playwright-against-previews, and the workflows
  ([ci.yml](../templates/github/workflows/ci.yml),
  [okr-gate.yml](../templates/github/workflows/okr-gate.yml),
  [neon-cleanup.yml](../templates/github/workflows/neon-cleanup.yml))
- branch protection via [branch-protection.sh](../templates/stack/branch-protection.sh)

Feature flags are absent from this list on purpose: flags are code, and the first one
installs itself — one package, one file, no setup step.

Greenfield installs reach it through `/okrdev:install`, which offers the stack module last —
after the method is in place, because the method is the product and the stack is a module.

SHA-256: 2dcd3922948d12fd186f99b53f5520eea8ea91108b7935420ad1ebaa23c7ae4a