Keystone
static-var v2.0.4
Publisher description
From the marketplace listing
Keystone helps developers move from ambiguity to verified delivery with 9 focused skills for understanding codebases, planning changes, implementing and refactoring safely, diagnosing failures, reviewing work, and shipping with evidence. Each skill keeps the current phase explicit, protects user-owned work, and requires proof before completion claims.
Language: English · Automatically detected from descriptions.
Publisher keywords
Search terms declared by the publisher.
Files & skills
File archives
Skill instructions
change-review13.2 KB
--- name: change-review description: Project change review for a concrete diff, branch, PR, patch, migration, fix, implementation result, or project plan that needs a new evidence-backed readiness verdict, or when another Keystone skill needs the Review Gate evaluated. --- # Change Review ## Core principle Change Review is an independent, read-only attempt to disprove readiness. Ask two questions at the same time: 1. **Spec axis:** does the work satisfy the stated requirements and acceptance criteria? 2. **Standards axis:** is it secure, correct, maintainable, tested, and safe to operate? Do not assume changed lines are the blast radius. Trace callers, callees, contracts, data flow, tests, runtime paths, and user impact before giving a verdict. ## Load when Load when a concrete software project artifact needs review: a diff, branch, PR, patch, migration, fix, implementation result, or project plan that must be assessed against its specification and engineering standards. Also load when another Keystone skill needs `../../references/gates/review.md` satisfied before shipping. At entry, use the full Keystone path when there is an inspectable project artifact and a readiness or blocker verdict is the outcome. Handle general critiques, prose review, and standalone explanations directly. Explicit invocation selects the full Change Review behavior. ## Not for Do not use Change Review for: - fixing, refactoring, formatting, or rewriting code - committing, merging, tagging, publishing, or shipping - initial implementation planning before a reviewable artifact exists - open-ended context-survey with no concrete artifact to assess - debugging where the requested outcome is a fix If asked to review and fix, review first, stop, and hand findings to `implementation`, `root-cause-analysis`, `context-survey`, `shipping`, or a human only after explicit permission. ## Outcome contract A complete review returns: - verdict: **Block**, **Caution**, or **Looks good** - findings ordered P0, P1, P2, P3, then Nitpicks - evidence for every finding: file/line, behavior path, contract, test, log, or doc - user impact and why the severity is justified - remediation guidance without applying the fix - tests that should be added or updated for affected behavior - scope reviewed, validation run, limitations, and read-only confirmation The review is incomplete if it only inspects the diff, only comments on style, or cannot explain how the work behaves at runtime. ## Change Review passes Perform multiple passes. New evidence from one pass expands later passes. ### Pass 0: scope and baseline - Identify artifact reviewed: diff, branch, files, release candidate, or plan result. - Read the user request, issue, spec, acceptance criteria, and claimed completion. - Check repository status without modifying files. - Record uncommitted work as context, not cleanup. ### Pass 1: spec compliance - Compare implementation against explicit requirements and non-goals. - Check edge cases, error states, and acceptance criteria. - Separate spec misses from standards concerns. - Treat a clean implementation of the wrong behavior as a finding. ### Pass 2: correctness and runtime paths - Trace primary success and failure paths end to end. - Follow changed functions into helpers, services, adapters, persistence, UI, jobs, and serializers. - Validate inputs, outputs, invariants, state transitions, retries, ordering, concurrency assumptions, and error propagation. - Look for nullability, off-by-one, time, encoding, pagination, caching, idempotency, cancellation, and partial-failure issues. ### Pass 3: regression and compatibility - Identify callers, consumers, and workflows that rely on old behavior. - Check public APIs, CLIs, schemas, migrations, persisted data, environment variables, feature flags, configuration defaults, and documentation. - Consider rollback, downgrade, mixed-version, and incremental rollout risks. - Search for tests or fixtures that encode previous behavior. ### Pass 4: security, privacy, and abuse resistance - Review authentication, authorization, tenancy, secrets, logging, validation, injection, XSS, SSRF, path traversal, unsafe deserialization, and RCE surfaces. - Check whether sensitive data leaks through errors, logs, telemetry, URLs, caches, exports, screenshots, or third-party calls. - Consider malicious users, compromised clients, replay, races, resource exhaustion, privilege escalation, and denial of service. ### Pass 5: tests and proof - Map changed behavior to existing tests. - Identify missing unit, integration, contract, regression, migration, security, accessibility, performance, or end-to-end coverage. - Prefer behavior assertions over implementation trivia. - Run focused read-only validation when practical: existing tests, type checks, lint, builds, or targeted commands. - If validation cannot run, state why and what should be run. ### Pass 6: maintainability and architecture - Assess clarity, cohesion, naming, dependency direction, duplication, complexity, observability, and debuggability. - Check architectural boundaries, local conventions, and API contracts. - Flag brittle abstractions, hidden coupling, unnecessary cleverness, and premature generalization when they create real maintenance risk. ### Pass 7: user impact and final consistency - Translate technical issues into affected personas, workflows, data, accessibility, performance, reliability, and support burden. - Re-rank findings by blast radius, likelihood, recoverability, and detectability. - De-duplicate findings, verify evidence, and state limitations honestly. ## Severity rubric Severity reflects realistic impact, not fix size. ### P0: Critical blocker Immediate or likely severe harm. Examples: - data loss, corruption, or irreversible destructive action - unauthorized access, privilege escalation, secret exposure, or major privacy breach - production outage or release artifact that cannot safely deploy - legal/compliance risk with material impact P0 means do not ship or merge without accountable human acceptance and mitigation. ### P1: Blocking defect High-impact issue that violates core requirements or creates serious regression risk. Examples: - primary workflow broken for a meaningful user segment - incorrect billing, permissions, persistence, or business logic - migration or compatibility gap that can break real deployments - high-risk behavior lacking tests plus a plausible failure mode P1 normally blocks shipping. ### P2: Important non-blocker or conditional blocker Material issue with bounded impact, lower likelihood, or workaround. Examples: - edge case with clear user impact - moderate-risk test gap - maintainability issue likely to cause near-term bugs - weak observability for a risky path State whether release context makes it blocking. ### P3: Low-risk improvement Valid concern with limited impact. Examples: - confusing name or local complexity that slows future work - minor non-hot-path performance inefficiency - incomplete docs for non-critical behavior - small test organization weakness P3 should not block unless it compounds with related risks. ### Nitpick Cosmetic, preference-level, or optional feedback: unenforced formatting, wording tweaks, or style suggestions with no correctness or maintainability impact. Keep nitpicks separate from severity findings. ## Impact tracing For each meaningful change, trace: - **Entry points:** user action, API route, CLI, job, event, hook, or import. - **Callers:** who invokes this and what assumptions they make. - **Callees:** helpers, libraries, persistence, network calls, and side effects. - **Data flow:** input, validation, transformation, storage, serialization, output. - **Contracts:** types, schemas, public APIs, flags, config, docs, and errors. - **Runtime paths:** success, failure, retry, timeout, cancellation, concurrency. - **Tests:** existing coverage, missing assertions, fixtures, mocks, snapshots. - **Users:** visible behavior, accessibility, performance, reliability, trust. If tracing leaves uncertainty, gather more read-only evidence or report the limitation. Do not invent confidence. ## Security and regression checklist Ask for every non-trivial review: - Can a user access, modify, infer, or delete data they should not? - Are authn, authz, tenancy, and ownership checked at the right layer? - Can untrusted input reach queries, interpreters, shells, paths, templates, redirects, or deserializers unsafely? - Are secrets, tokens, PII, or internal identifiers exposed in logs, errors, telemetry, URLs, caches, or client bundles? - Did defaults, permissions, feature flags, or safeguards become unsafe? - Are races, duplicate submissions, retries, replay, and out-of-order events safe? - Can persisted data be corrupted, stranded, or made hard to rollback? - Are public APIs, stored data, configs, and integrations backward compatible? - Does failure degrade safely without hidden partial success? - Are performance, resource use, accessibility, localization, and platform differences acceptable for realistic users and abuse? - Do tests cover the affected behavior and important regression paths? ## Subagents and reasoning Use read-only subagents for separable risks: security/privacy, test coverage, architecture/API compatibility, persistence/migration, accessibility/user impact, performance, concurrency, or release risk. Use deeper analysis for security-sensitive, data-loss, billing, permissions, public API, migration, or cross-system reviews. When delegation is available, encode required evidence depth and review standard in the prompt. Subagents must receive the read-only contract and return evidence-backed findings, not patches. Reconcile duplicates and conflicts before reporting. The primary reviewer owns final severity and verdict. ## Hard rules - Read-only: inspect files without editing, formatting, generating, staging, committing, merging, tagging, or publishing them. - Report issues without silently fixing them. - Do not run destructive or project-mutating commands. - Do not rely only on changed lines; inspect impacted code paths and contracts. - Do not approve solely because tests pass. - Do not report speculation as fact; mark uncertainty. - Do not bury blockers under minor comments. - Do not disguise style preferences as correctness findings. - Do not omit needed tests when behavior changed. - Do not satisfy `../../references/gates/review.md` unless blockers and non-blockers are separated. - Load `../../references/gates/review.md` before the verdict; Change Review owns the review execution and supplies the gate evidence. - Run the checkpoint gate before the final response; if review passes and delivery/finalization was requested, route to `shipping` or leave an explicit shipping prompt. ## Failure modes Avoid these anti-patterns: - **Single-pass skim:** one read of changed lines plus generic comments. - **Diff tunnel vision:** missing callers, callees, contracts, and user impact. - **Checklist theater:** naming security/tests without tracing actual risk. - **Green-test rubber stamp:** assuming current tests prove new behavior. - **Spec blindness:** judging code quality while requirements are unmet. - **Standards blindness:** accepting unsafe or fragile code because the narrow spec passes. - **Severity inflation:** turning preferences into blockers. - **Severity deflation:** downgrading real user harm because the fix is small. - **Patch creep:** fixing, refactoring, or committing instead of reviewing. - **Unowned uncertainty:** failing to state what was not verified. - **Lost next event:** Change Review passes but never routes or prompts for `shipping` when finalization remains. ## Output format Worked finding example: ```markdown ### P1 - Missing tenant check on invoice export - Evidence: `api/exportInvoice.ts:42` accepts `invoiceId` and loads the invoice without comparing `invoice.accountId` to the authenticated account; `/invoices/:id/export` is reachable by any logged-in user. - Impact: A user who guesses another invoice ID can download billing data from a different account, which is a privacy and authorization breach. - Recommendation: Enforce tenant ownership before export and return the existing unauthorized response on mismatch. - Tests needed: Add an integration test where account A requests account B's invoice and receives 403/no file, plus a happy-path same-account export test. ``` Use this structure: ```markdown ## Verdict Block | Caution | Looks good ## Scope reviewed - Artifact reviewed: - Key files/paths inspected: - Validation run: - Review limitations: - Read-only confirmation: no files changed by this review ## Findings ### P0 - [Title] - Evidence: - Impact: - Recommendation: - Tests needed: ### P1 None ### P2 None ### P3 None ## Nitpicks None ## Tests to add or update - Behavior: - Suggested coverage: - Why it matters: ## Handoff - Blockers: - Non-blocking follow-up: - Suggested owner module: implementation, root-cause-analysis, context-survey, shipping, or human ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ``` If a severity has no findings, write `None`. Recommendations must be actionable but must not be applied by Change Review. ## Shared standards For architecture-sensitive or code-quality-sensitive work, load `../../references/engineering-standards.md` and apply it as reference, not dogma.
context-survey6.16 KB
--- name: context-survey description: Project context survey for repository reconnaissance or project decisions that require evidence from existing code, project documentation, or project-specific external sources before planning or changing the project. --- # Context Survey ## Core principle Context Survey is evidence gathering before action. Inspect available material first, preserve source quality, separate facts from assumptions, and stay read-only unless the user explicitly asks for a durable context-survey artifact. ## Load when Load when a software or product project needs: - repository reconnaissance or understanding as the requested project outcome; - evidence from existing code, tests, configuration, history, or project documentation; - project-specific external documentation, standards, issues, market examples, or APIs; - synthesis or validation of project evidence before another Keystone phase. ## Entry fit An explicit invocation selects the full Context Survey behavior. A Keystone handoff selects the full skill for project-bound repository understanding or evidence needed for a project decision or change. For automatic entry, the prompt's subject must be project-bound. Load the full skill when repository understanding or reconnaissance is the requested outcome, or when this evidence contract materially improves a project decision or change. The current directory alone does not establish project scope. For automatic routing of standalone public search, general research, ordinary explanation, or summarization unrelated to a project, use the direct research/answering path and return the requested result without Keystone gates. ## Not for - Implementing, refactoring, editing, or fixing code. - Shaping product direction beyond evidence-backed options. - Broad tooling risk audits; use `project-audit`. - Root-cause repair of a failure; use `root-cause-analysis` after initial context. - Guessing when evidence can be inspected. ## Outcome contract Deliver a context-survey brief that states: - question or decision being supported; - sources inspected, with file paths, commands, URLs, or other citations; - source-quality notes (primary vs secondary, current vs stale, authoritative vs anecdotal); - findings separated from assumptions and unknowns; - confidence level and why; - recommended next module, or `none` if no Keystone handoff is warranted. ## Modes - **Repository read:** inspect files, history, configs, tests, docs, and existing behavior. Prefer primary project evidence. - **External context-survey:** compare outside documentation, standards, issues, market examples, or APIs. Cite URLs and note recency. - **Synthesis:** combine several sources into a decision-ready summary with tradeoffs and confidence. - **Discovery scout:** map a large unknown area without drawing strong conclusions until evidence is sampled. ## Process 1. Restate the context-survey question or requested repository reconnaissance outcome, and the downstream decision when applicable. 2. Inspect before asking: search/read the repo, docs, logs, or provided sources before requesting more context. 3. Prefer primary evidence: source code, tests, product docs, official docs, reproducible commands, direct user-provided material. 4. Track citations as you go. Every important claim should point to evidence or be labeled as an assumption. 5. Evaluate source quality: age, authority, completeness, bias, and whether evidence is direct or inferred. 6. Compare alternatives when relevant, including costs, risks, constraints, and implications of taking no action. 7. State unknowns explicitly. Do not fill gaps with confident-sounding speculation. 8. Recommend the smallest next step: `product-planning`, `root-cause-analysis`, `project-audit`, `task-creation`, `implementation`, `change-review`, or stop. ## Subagents and reasoning Use read-only subagents when the search space is large or evidence can be gathered independently. Use lightweight analysis for narrow file summaries and deeper analysis when findings affect architecture, security, safety, release decisions, legal/market claims, or irreversible product direction. When delegation is available, encode required evidence depth and risk standard in the prompt. Subagents must remain read-only unless the user requested an artifact. ## Hard rules - No mutation by default: do not edit files, run formatters, or alter state except harmless read-only commands. - Durable artifact exception: if the user requests a context-survey artifact, confirm the path and scope before writing; hand off to `implementation` when the artifact changes product/code behavior or touches broader project structure. - Cite evidence for material claims; if evidence is unavailable, say so. - Distinguish facts, interpretations, assumptions, and recommendations. - Do not ask for information that can be inspected first. - Do not present search results or model knowledge as authoritative without source-quality caveats. ## Failure modes - **Context theater:** long summaries without citations or decision relevance. - **Source laundering:** treating blogs, stale docs, or guesses as facts. - **Premature shaping:** deciding product behavior before evidence is clear. - **Mutation creep:** “just fixing” or rewriting while context-surveying. - **Hidden uncertainty:** omitting confidence, unknowns, or contradictory evidence. ## Worked example Good context-survey finding: “Official Stripe docs show idempotency keys apply per unique key and preserve the first result, including failures; this means retrying payment capture should reuse the original key, not generate a new one. Confidence: High — primary docs, current page.” Bad context-survey finding: “Stripe probably handles retries safely, so we can just retry the request.” ## Output format ```markdown ## Context Survey brief Question: ... ### Evidence inspected - `path/or/source`: what it shows, quality note ### Findings - Fact — citation - Interpretation — citation + reasoning ### Assumptions / unknowns - ... ### Options or implications - ... ### Confidence High/Medium/Low — why ### Recommended next step Module or `none`, with rationale ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ```
implementation14.8 KB
---
name: implementation
description: Executable-software implementation in an identified software project for features, diagnosed fixes, migrations, integrations, and build or release automation. Select only when the requested change alters runtime, API, data, build, or release behavior and requires project-specific verification of integrated effects.
---
# Implementation
## Core principle
Implementation changes through evidence, not vibes: isolate first, specify the next observable behavior, prove the test can fail, make the smallest correct change, then refactor without changing behavior.
Implementation is the mutation module. It may edit scoped project artifacts after the isolation gate passes, but it does not decide that work is shipped. Completion means "implemented with proof and handed to a change-review/shipping checkpoint," not finalized.
## Load when
Use Implementation when a named software or digital-product initiative needs:
- a feature, screen, endpoint, command, migration, or integration implemented;
- a diagnosed fix applied;
- architecture, module boundaries, interfaces, or contracts changed;
- an approved plan from `task-creation` executed;
- independent implementation slices delegated when a subagent tool is available;
- supporting artifacts only within an already-active behavior-changing implementation: required project configuration, documentation, content, or scripts.
## Entry fit
An explicit invocation selects the full Implementation behavior.
A Keystone handoff selects the full skill when a project change needs this mutation-and-proof contract.
For automatic entry, load the full skill when the prompt's subject is project-bound, the requested change alters executable software behavior, **and** this mutation-and-proof contract materially improves the change. The current directory alone does not establish project scope.
For automatic routing of a mechanical single-file configuration edit, one-off script, or other standalone small mutation that does not need project-specific context or proof gates, use the direct editing path and verify the requested result proportionately.
## Not for
Do not use Implementation for:
- routing unclear work: use `context-survey`
- context-surveying unknowns without mutation: use `context-survey`
- shaping requirements or acceptance criteria: use `product-planning`
- decomposing large work before implementation: use `task-creation`
- diagnosing a failure whose cause is unknown: use `root-cause-analysis`
- reviewing completed changes: use `change-review`
- releasing, merging, publishing, or finalizing: use `shipping`
## Outcome contract
Before Implementation exits, it must be able to report:
- isolation was checked before the first mutation via `../../references/gates/isolation.md`
- the exact user scope and protected files were respected
- the intended behavior or refactor invariant is stated plainly
- tests, examples, or checks prove the change, or gaps are explicitly disclosed
- `../../references/gates/red.md` passed for behavior changes
- `../../references/gates/proof.md` passed before any success claim
- delegated work, if any, was verified by the parent before acceptance
- `../../references/gates/checkpoint.md` decided the next required event
Implementation must not claim work is done because code "looks right." Proof and an explicit change-review/shipping checkpoint are required before completion claims.
## Modes
### TDD feature implementation
Use when adding or changing observable behavior.
Contract:
- define one behavior slice at a time
- write or identify the smallest test/check that should fail before the change
- run it and confirm the failure is meaningful, not caused by setup noise
- implement the smallest code that makes it pass
- run the focused check again
- run relevant regression checks
- refactor only while checks stay green
Load and pass `../../references/gates/red.md` for the red signal or its explicit exception. Prefer checks that exercise real behavior over mocks, implementation details, or snapshots.
### In-slice refactor
Use for behavior-preserving cleanup within an already-active behavior-changing implementation, or after an explicit Keystone handoff from `refactoring`. Route standalone behavior-preserving structural work to `refactoring`.
Contract:
- name the behavior that must not change
- find existing tests or add characterization tests before edits when coverage is weak
- make small mechanical changes first
- keep public contracts stable unless the user requested a contract change
- run regression checks before and after meaningful refactor steps
- stop if behavior questions appear; route back to `product-planning` or `root-cause-analysis` as needed
A characterization test captures what the current system does before you change structure. It is not a claim that current behavior is ideal; it is a tripwire that prevents accidental behavior changes while refactoring. Write it around externally visible behavior, important edge cases, or bug-compatible outputs that must stay stable until the user approves a behavior change.
Concise example: before extracting invoice total formatting, add a test that `renderInvoiceSummary({ subtotal: 1000, discount: 125, currency: "USD" })` still returns `"Subtotal $10.00 · Discount $1.25 · Total $8.75"`; then refactor behind that externally visible output.
Refactoring is not a license to redesign everything nearby.
### Architecture-sensitive implementation
Use when the change affects boundaries, state management, cross-module dependencies, platform conventions, or long-lived maintainability.
Contract:
- identify the domain shape before choosing architecture
- choose the lightest architecture that fits the problem
- define contracts/interfaces only at real boundaries where they reduce coupling or clarify ownership
- keep domain rules separate from transport, UI, persistence, and framework glue
- use patterns only when they remove current pressure, not to decorate simple code
- validate observable maintainability with the pressure-test below
Architecture pressure-test:
- Does the architecture match domain complexity rather than a desire to sound senior?
- Is each module/class/function responsible for one understandable thing?
- Are domain concepts named in the user's language?
- Is state ownership explicit, with side effects isolated enough to test meaningful behavior?
- Would a developer know where to add the next similar behavior?
- Is the simplest path readable, or has simplicity become cleverness?
- Is duplication removed only after the repeated concept is real?
- Do patterns such as Strategy, Repository, Adapter, Observer, MVVM, MVI, or Clean Architecture relieve current pressure?
- **SOLID check:** one reason to change; known variation handled safely; substitutes honor contracts; callers avoid unused surface; high-level policies depend on stable abstractions only at real boundaries.
- Smell stop-list: god functions, vague managers/helpers, hidden control flow, stringly APIs, layer violations, speculative interfaces.
### Delegated/parallel implementation
Use when two or more implementation slices are independent enough to proceed without shared mutable state or ambiguous ownership.
Contract:
- split by outcome, not by vague activity
- pass each worker a narrow scope, files or directories, constraints, and completion criteria
- define contracts/interfaces before parallel work begins when slices must meet
- state what must not be edited
- require each worker to report files changed, checks run, and risks
- verify delegated work yourself before integrating or claiming completion
- reconcile overlaps deliberately; do not let workers race on the same files
Do not delegate fuzzy architecture judgment without a concrete interface or decision boundary.
Concise subagent brief template:
```markdown
Goal: one observable outcome
Scope: allowed files/directories
Do not edit: protected files/behaviors
Contract: interfaces, invariants, data shape, or acceptance criteria
Proof: required tests/checks/manual verification
Report: files changed, verification output, risks/gaps
```
## Process
1. Confirm scope.
- Restate the requested mutation in one sentence.
- Identify protected files and out-of-scope behavior.
- If scope is unsafe or ambiguous, ask one focused question or route to `product-planning`.
2. Pass isolation before mutation.
- Load/check `../../references/gates/isolation.md`.
- Know the workspace, branch/worktree state, and dirty files.
- Stop if unrelated changes could be overwritten.
3. Choose the mode.
4. Define the next outcome.
- Write the smallest observable behavior, invariant, or contract.
- Avoid "make it better" as an implementation target.
5. Establish the signal before code.
- For behavior changes, load and pass `../../references/gates/red.md` before editing.
- For refactors, establish characterization or regression coverage.
6. Implement the smallest correct slice.
- Edit only files in scope.
- Keep changes narrow and reversible.
- Prefer clear names and direct control flow.
- Do not invent broad architecture to satisfy a small behavior.
7. Green.
- Run the focused test/check.
- If it fails unexpectedly, use `root-cause-analysis`; do not stack guesses.
8. Refactor.
- Remove duplication introduced by the slice.
- Improve readability without broadening behavior.
- Keep tests green after cleanup.
9. Regression check.
- Load and pass `../../references/gates/proof.md` against the intended outcome.
10. Early smell check.
- Stop and simplify if the diff hits the architecture smell stop-list.
11. Architecture pressure-test.
- For architecture-sensitive changes, answer the inline pressure-test and remove abstractions that do not survive it.
12. Checkpoint and handoff.
- Load/check `../../references/gates/checkpoint.md`.
- Summarize changed files and behavior.
- Include commands run and results.
- Disclose unverified areas.
- Decide whether `change-review` is required now, can be satisfied by self-review, or must be left as a pending review pointer.
- If review is required and Keystone can safely continue, hand off to `change-review` before the final response. If not, ask the user or include the pending review pointer from `../../references/gates/review.md`.
## Subagents and reasoning
Use deeper analysis when architecture boundaries are being changed, tests fail for unclear reasons, multiple agents must coordinate through contracts, data loss/security/billing/release/migration risk exists, or the user asks for broad refactoring or platform-specific architecture judgment. When delegation is available, encode required evidence depth and risk standard in the prompt.
Use read-only exploration subagents when independent inspection can reduce uncertainty without mutation. Use implementation subagents only when the user requested mutation, isolation has passed, allowed files are scoped, each slice can be verified independently, and file overlaps are prevented by ownership or explicit handoff.
A delegation brief must include:
- goal/outcome
- allowed files or directories
- forbidden files or behaviors
- relevant interfaces/contracts
- expected tests/checks
- report format for files changed, verification, and risks/gaps
Verify delegated work by:
- reading the diff, not just the report
- running or reviewing the reported checks
- checking contract compatibility between slices
- rejecting broad edits, invented abstractions, or unproved claims
## Hard rules
- Implementation must pass `../../references/gates/isolation.md` before the first mutation.
- Implementation must not ship, merge, publish, release, or finalize work.
- Implementation must run the checkpoint gate before any final response.
- After mutation, Implementation must not stop at “implemented” when review remains; continue to `change-review` when safe or leave an explicit review prompt/pending pointer.
- Do not edit files outside the user's scope.
- Do not claim completion without proof or explicit disclosure of missing proof.
- Do not skip red/green/refactor for behavior changes unless there is a stated, practical reason and alternative proof plan.
- Do not use subagents as a way to avoid understanding the result.
- Do not introduce architecture that the current domain pressure does not justify.
## Failure modes
Watch for these and correct course:
- **Green-only development:** tests are added after implementation and never proven red.
- **Mock theater:** tests prove mocks were called, not that behavior works.
- **Scope creep:** nearby cleanup, README edits, scripts, or unrelated modules change without request.
- **Invented architecture:** factories, managers, providers, repositories, or layers appear without pressure.
- **Shallow abstraction:** a wrapper hides one call site and makes the code harder to follow.
- **Brittle hack:** timing sleeps, magic constants, global state, or special cases mask the real issue.
- **God function:** one routine owns validation, orchestration, persistence, formatting, and error policy.
- **Vague names:** `manager`, `helper`, `util`, or `common` hide responsibility instead of naming it.
- **Hidden control flow:** callbacks, observers, magic registration, or framework hooks make execution hard to trace without clear benefit.
- **Stringly API:** strings encode commands, states, fields, or permissions that should be typed, enumerated, or centralized.
- **Layer violation:** UI, transport, persistence, or framework code reaches across boundaries into another layer's policy.
- **Speculative interface:** abstraction exists for imagined future variants, not current pressure.
- **Ambiguous delegation:** workers receive goals like "improve this" without contracts or boundaries.
- **Unverified handoff:** delegated changes are accepted from a summary alone.
- **Finalization leak:** Implementation says work is shipped, merged, or ready for users instead of ready for change-review/shipping.
- **Lost next event:** Implementation ends with “done” or only a passive Next line while change-review, shipping, or a user approval is still required.
## Output format
When handing back from Implementation, respond with these sections, including `### Checkpoint`:
- `Summary`: what changed, in bullets
- `Files changed`: exact paths
- `Verification`: commands/checks run and results
- `Delegation`: subagents used, contracts passed, and how their work was verified; or `none`
- `Risks / gaps`: anything unverified, deferred, or worth reviewing
- `### Checkpoint`: current skill, completed gates, next required skill, next check, action (`continue now`, `ask user`, `pending pointer`, or `stop`)
- `Next`: usually `change-review` or `shipping`, phrased as an action/prompt or already-executed handoff; never a claim that Implementation finalized the work
## Shared standards
For architecture-sensitive or code-quality-sensitive work, load `../../references/engineering-standards.md` and apply it as reference, not dogma.
product-planning10.2 KB
--- name: product-planning description: Software product planning for a concrete software or digital-product initiative whose behavior, UX, user-facing copy, technical direction, scope, or acceptance criteria must be shaped before implementation, or whose approved direction must become a specification. --- # Product Planning ## Core principle Product Planning is a specification algorithm: turn an unclear intent into exact behavior, constraints, tradeoffs, and acceptance criteria before anyone implements. It decides what should be true, not whether code is complete. ## Load when Load when a concrete software or digital-product initiative needs: - product behavior, business rules, scope, success criteria, or acceptance criteria shaped before implementation; - UX/UI flows, states, accessibility behavior, or visual constraints decided; - user-facing product copy shaped as part of a feature or product experience; - technical direction, boundaries, APIs, data flow, or architecture tradeoffs decided; - alternatives explored for a project decision; - approved direction converted into a specification; - context-survey evidence converted into an implementation-ready direction. ## Entry fit An explicit invocation selects the full Product Planning behavior. A Keystone handoff selects the full skill when a concrete software or digital-product initiative needs this specification contract. For automatic entry, load the full skill when the prompt's subject identifies a concrete software or digital-product initiative **and** this specification contract materially improves the decision. The current directory alone does not establish project scope. For automatic routing of ordinary brainstorming, standalone copywriting, rewriting, naming, or visual ideation unrelated to a concrete software or digital-product initiative, use the direct creative/answering path and return the requested result without Keystone gates. ## Not for - Writing implementation code or changing runtime behavior; hand off to `implementation`. - Diagnosing failures; use `root-cause-analysis`. - Broad repository audit or release readiness; use `project-audit` or `shipping`. - Inventing facts that should be context-surveyed first. - Polishing completed work as shippable proof. ## Outcome contract Deliver a shaped proposal that includes: - goal, user/audience, and success criteria; - product behavior and UX states, including happy, empty, loading/pending, error/failure, and edge/constraint states where relevant; - copy or content direction when user-facing text matters; - architecture and scope tradeoffs at the level needed for planning, not implementation; - alternatives considered and why one direction is preferred; - acceptance criteria and non-goals; - recommended next module (`task-creation`, `implementation`, `change-review`, `context-survey`) or `none` if no Keystone handoff is warranted. ## Modes - **Product planning:** specify the user job, trigger, actor permissions, core flow, business rules, constraints, success metrics, non-goals, and acceptance criteria. - **UX/UI planning:** specify layout hierarchy, navigation, interaction model, responsiveness, accessibility, visual constraints, and the 5-state UX checklist: happy, empty, loading/pending, error/failure, edge/constraint. - **Copy planning:** specify audience, message hierarchy, claims, tone, CTA, labels, empty/error text, and prohibited vague claims. - **Technical planning:** specify boundary placement, API granularity, data flow, state ownership, dependency direction, persistence/integration seams, and architectural tradeoffs without writing code. - **Alternative exploration:** present multiple viable directions before choosing or asking the user to choose. ## Process 1. Classify the request into one or more modes: product, UX/UI, copy, technical, or alternatives. 2. Identify the goal, primary user/audience, job-to-be-done, context of use, and success criteria. 3. Inspect existing product patterns, domain language, and provided material before inventing new conventions. 4. Convert intent into exact rules: - Product: actor, trigger, preconditions, action, result, permissions, limits, and measurable success. - UX/UI: screen/region hierarchy, controls, transitions, accessibility behavior, responsive behavior, and 5-state UX checklist. - Copy: exact headline/body/CTA/error text or content rules, with claims grounded in known facts. - Technical: components/modules involved, ownership boundaries, contracts, data flow, failure handling, migration or rollout constraints. 5. Apply technical shaping heuristics when architecture matters: - **Boundary placement:** put boundaries where ownership, volatility, testability, or external systems change; do not split stable one-step logic. - **API granularity:** prefer operations that match caller intent; avoid both chatty micro-methods and god endpoints that hide unrelated behavior. - **Data flow:** name source of truth, state transitions, sync/async edges, validation points, and where errors surface. - **Architectural tradeoffs:** state what becomes simpler, harder, slower, safer, more testable, or more coupled. 6. Ban fluffy terms unless translated to behavior. Words like “modern,” “clean,” “intuitive,” “delightful,” “seamless,” or “user-friendly” must become observable rules. 7. Offer alternatives when the direction is not obvious. Include the “do nothing / decide later” option if legitimate. 8. Convert the chosen direction into acceptance criteria that can be implemented and reviewed. 9. Stop at the spec boundary. If the user asks for design plus implementation, finish Product Planning with the spec and recommended handoff to `implementation`; do not implement code. ## Subagents and reasoning Use subagents for bounded alternatives, critique, or parallel concepts when the active host exposes safe delegation. Use lightweight analysis for narrow copy/behavior edits and deeper analysis for multi-screen flows, accessibility-sensitive experiences, design-system impact, pricing/positioning, architecture boundaries, or major scope decisions. When delegation is available, encode required evidence depth, constraints, and risk standard in the prompt. Subagents should produce options or critique, not unrequested implementation. ## Hard rules - Product Planning is not implementation: do not edit production code or runtime behavior. - If the user asks for design and implementation together, Product Planning stops after the specification and hands off to `implementation`. - Ground claims in context-survey or existing product evidence; call `context-survey` when facts are missing. - Always identify user/audience and success criteria for product-facing work. - Include acceptance criteria before handing off to implementation. - Translate fluffy descriptors into exact behavior; otherwise remove them. - Avoid action bias: if the best answer is “do nothing” or “decide later,” say so with criteria. ## Failure modes - **Abstract advice:** principles without actors, states, rules, tradeoffs, or acceptance criteria. - **Pretty but unusable:** visual ideas without behavior, states, or acceptance criteria. - **Fluffy spec:** “modern/user-friendly” language without exact behavior. - **Spec as proof:** implying a design solves the problem before implementation or validation. - **Audience blur:** writing for everyone and satisfying no one. - **Scope fog:** hiding hard tradeoffs until implementation time. - **Premature code:** implementing while still deciding what should exist. ## Examples Good product planning: “When a workspace has no projects, show an empty state with title ‘Create your first project,’ one-sentence explanation, primary ‘New project’ CTA, and no table chrome. Success: first project creation rate increases.” Bad product planning: “Make the dashboard more useful and modern.” Good UX/UI planning: “On save, disable the Save button, keep the form editable fields visible, show inline progress text ‘Saving…’, then restore focus to the first invalid field on failure.” Bad UX/UI planning: “Use a clean, user-friendly save experience.” Good copy planning: “CTA says ‘Start free trial’ because billing is not required; avoid ‘Buy now.’ Error text names the failed action and recovery: ‘We couldn’t send the invite. Check the email address and try again.’” Bad copy planning: “Use friendly copy that reduces friction.” Good technical planning: “Keep validation in the domain service because API and background import both need it; expose one `createInvite` operation that returns accepted, duplicate, or invalid-email outcomes.” Bad technical planning: “Add a helper/manager layer so the architecture is scalable.” Worked technical planning: “For export retries, keep the queue worker as the owner of retry state, expose `requestExport(accountId, format)` from the API, persist `pending|running|failed|ready` status in `exports`, and surface failures through the existing job status endpoint. Tradeoff: one extra status read, but retry policy stays out of controllers and can be tested without HTTP.” ## Output format Always include goal, audience/user when relevant, mode(s), acceptance criteria, and recommended next step. Include UX/copy/technical/alternatives sections only when they change the decision; omit empty or irrelevant headings. ```markdown ## Planned direction Goal: ... Audience/user: ... Mode(s): product | UX/UI | copy | technical | alternatives ### Proposed behavior / experience - Actor/trigger/preconditions: ... - Rules/results: ... ### UX states and copy (when relevant) - Happy: ... - Empty: ... - Loading/pending: ... - Error/failure: ... - Edge/constraint: ... - Key copy: ... ### Technical planning (when relevant) - Boundaries/API/data flow: ... - Tradeoffs: ... ### Scope and tradeoffs - In: ... - Out: ... - Tradeoffs: ... ### Alternatives considered (when relevant) - Option A: ... - Option B / do nothing: ... ### Acceptance criteria - ... ### Recommended next step Module or `none`, with rationale ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ``` ## Spec artifact rule Create a spec file only after the user has finalized or approved the plan details. Before approval, work in conversation. Default path: `docs/keystone/specs/YYYY-MM-DD-<slug>.md`.
project-audit11.9 KB
--- name: project-audit description: Project health audit for a concrete software repository or product subsystem. Use when its tooling, CI, dependencies, configuration, tests, documentation, packaging, architecture, or agent instructions need a whole-system condition and risk assessment. --- # Project Audit ## Core principle Project Audit is a read-only, whole-project condition scan. It turns repository evidence into mechanical findings about system drift, fragility, and maintenance risk; it does not repair, refactor, update, or reconfigure anything by default. Project Audit answers “what is the condition of this project or subsystem?” Change Review answers “is this specific change acceptable?” If the task centers on a PR/diff/patch, use `change-review`; if it centers on repo-wide readiness, drift, tooling, docs, tests, releases, or operational condition, use `project-audit`. ## Load when Load when the user asks for a whole-system condition scan of a concrete software repository or product subsystem: project health, release readiness, tooling/CI drift, dependency/config condition, maintenance risk, architecture smells, docs-vs-reality, test health, packaging, or instruction drift. At entry, use the full Keystone path when repository or subsystem evidence must support a project-level health verdict. Handle general risk questions, personal health checks, and narrow standalone inspections directly. Explicit invocation selects the full Project Audit behavior. ## Not for - Fixing issues found during the audit. - Reviewing a specific change, PR, patch, or commit; use `change-review`. - Tracing a specific failure to root cause; use `root-cause-analysis`. - Shipping a completed branch; use `shipping`. - Designing new product behavior; use `product-planning`. - Context Survey briefs about a narrow question; use `context-survey`. ## Outcome contract Deliver a project audit report where every finding maps: `finding -> evidence -> impact -> confidence -> next Keystone module` The report must include: - audit scope and evidence inspected; - concrete checklist status for tooling/CI drift, docs-vs-reality, config/env rot, dependency health, test health/flakiness, package/release health, and instruction/skill drift; - risks ranked by severity and confidence using the Project Audit priority rubric; - checks not run and why; - explicit no-fix confirmation unless repairs were requested. Project Audit priority rubric: - **Critical:** Broken now, release-blocking, security-sensitive, data-loss-prone, or prevents required project operation. Urgency: act before `shipping` or before depending on the affected subsystem. Next Keystone module: usually `root-cause-analysis` for failing behavior or `implementation` for repairs; use `shipping` only for release gate follow-up after the issue is fixed. - **Watch:** Risky, stale, drifting, or likely to become blocking, but not proven broken under current evidence. Urgency: schedule remediation or investigation soon; do not let it become untracked backlog. Next Keystone module: usually `context-survey` to verify unknowns, `task-creation` to plan multi-step remediation, or `implementation` for contained repairs. - **Info:** Health signal, minor inconsistency, low-impact cleanup, or explicitly unknown/unchecked area. Urgency: no immediate action unless priorities change. Next Keystone module: `none` when informational, or `change-review`/`shipping` when the next step is validation rather than repair. ## Modes - **Project snapshot:** summarize structure, active areas, scripts, tests, CI, docs, and current branch state. - **Tooling/CI drift audit:** compare manifests, scripts, Make targets, task runners, hooks, CI workflows, matrix versions, cache keys, and required checks against actual files and commands. - **Docs-vs-reality audit:** compare README, contributor docs, runbooks, release docs, examples, and command snippets with actual repo layout and executable scripts. - **Config/env rot audit:** inspect templates, `.env.example`, config schemas, secrets documentation, Docker/compose/devcontainer files, SDK/runtime pins, and environment assumptions for staleness or inconsistency. - **Dependency health audit:** inspect manifests, lockfiles, version pins, deprecated packages, engine/toolchain constraints, known update pressure, and manifest-lock consistency. - **Test health/flakiness audit:** inspect test commands, skipped/quarantined tests, retries, snapshots, coverage signals, slow/flaky markers, CI-only behavior, and failure history when available. - **Package/release health audit:** inspect package metadata, allowlists, build artifacts, version/changelog flow, release scripts, publish dry-run support, tags, and multi-target release paths. - **Instruction/skill drift audit:** inspect project instructions, agent docs, skill/module docs, validator rules, and examples for contradictions or stale references. - **Release readiness audit:** assess project-level risk categories before `shipping`, without preparing release artifacts. - **Risk triage:** rank issues by impact, likelihood, evidence strength, and recommended next module. ## Process 1. Define scope: whole repo, subsystem, tooling, release readiness, dependencies, docs, instructions, or operational process. 2. Confirm read-only posture. Say what will be inspected and avoid state-changing commands unless the user explicitly requested repairs. 3. Inspect before judging: read manifests, scripts, CI, tests, docs, configs, package/release files, instruction files, recent status, and relevant gates. 4. Apply the concrete audit checklists: - **Tooling/CI drift:** declared scripts exist; CI calls valid commands; local and CI tool versions align; required checks match current project; generated/cache paths are current; hooks and Make/package targets agree. - **Docs-vs-reality:** documented setup/test/implementation/release commands exist; referenced paths still exist; examples match current APIs/CLIs; screenshots/output snippets are not misleading; contributor docs match workflow. - **Config/env rot:** sample env files cover required variables; config names match code; defaults are safe; obsolete variables are not documented as required; runtime/container/dev environment pins are current. - **Dependency health:** manifests and lockfiles agree; package managers are not mixed accidentally; runtime engine constraints are plausible; deprecated/abandoned/high-risk dependencies are called out with evidence; update risk is separated from breakage. - **Test health/flakiness:** test entrypoints are discoverable; skips/todos/quarantine/retry settings are listed; flaky markers or timing-sensitive tests are noted; coverage signals are reported only if measured; CI-only gaps are identified. - **Package/release health:** package allowlists include required files and exclude junk; build artifacts are reproducible or documented; version/changelog/release scripts align; publish dry-run or validation exists; multi-target outputs are accounted for. - **Instruction/skill drift:** AGENTS/CLAUDE/GEMINI/Codex/plugin docs and Keystone skill docs agree; module boundaries are current; validators and examples match required headings/behavior. 5. Run safe focused checks when useful and allowed, such as `git status --short`, listing workflow files, reading manifests, `--help`, dry-run validation, or project validators. Avoid install/update/format/fix/publish commands by default. 6. Detect drift mechanically: documentation pointing to missing scripts, scripts referencing missing files, stale generated assets, inconsistent versions, orphaned configs, CI mismatch, package metadata mismatch, or contradictory instructions. 7. For each finding, write the required chain: severity, finding, evidence, impact, confidence, and next Keystone module (`context-survey`, `root-cause-analysis`, `product-planning`, `task-creation`, `implementation`, `change-review`, `shipping`, or `none`). Assign severity from the Project Audit priority rubric; do not produce unranked findings. Use `none` only when no Keystone handoff is warranted. If the finding is about tests or package metadata, still route to a real module, usually `implementation` for repairs, `root-cause-analysis` for failing behavior, `change-review` for validation, or `shipping` for final package readiness. 8. Separate statuses: **broken now**, **risky**, **stale**, **unknown**, and **healthy**. Map broken-now items to Critical unless evidence shows low impact; map risky or stale items to Watch unless release/security impact makes them Critical; map healthy, minor, and unchecked informational notes to Info. Do not convert unknowns into failures. 9. Stop at reporting unless the user explicitly requested fixes. If repairs are requested, route to the appropriate module instead of silently switching modes. ## Subagents and reasoning Use read-only subagents for broad inventory or independent risk triage when the active host exposes safe delegation. Use deeper analysis for release readiness, security-sensitive audits, large monorepos, severe tooling drift, instruction drift affecting agent behavior, or when audit findings affect go/no-go decisions. When delegation is available, encode required evidence depth and risk standard in the prompt. Subagents must remain read-only unless repairs are explicitly requested. ## Hard rules - Read-only by default: no fixing, formatting, dependency updates, cleanup, generation, or config changes unless explicitly requested. - Project Audit is whole-project/system condition; Change Review is a specific change. Do not use Project Audit to approve a PR diff. - Every finding must include severity, finding, evidence, impact, confidence, and next Keystone module. - Use only the Project Audit priority rubric for severity (`Critical`, `Watch`, `Info`) unless the user explicitly requests equivalent labels; do not emit unranked dumps. - Evidence categories must be named; unchecked areas must be listed with reasons. - Do not overstate confidence. Label inferred risks and explain what would verify them. - Prefer safe read-only or focused validation commands. - Separate “broken now,” “risky,” “stale,” and “unknown.” - Project Audit can recommend `shipping`, but does not replace shipping proof gates. ## Failure modes - **Audit-as-fix:** making opportunistic changes during a scan. - **Change Review confusion:** judging a specific PR/change instead of system condition. - **Checklist theater:** listing categories without evidence. - **Evidence gaps:** reporting findings that lack impact, confidence, or next module. - **False certainty:** declaring healthy because a narrow check passed. - **Drift blindness:** missing mismatches between docs, scripts, CI, package metadata, instructions, and actual files. - **Unranked dump:** overwhelming the user with findings but no severity or next module. ## Output format ```markdown ## Project Audit report Scope: ... No-fix status: confirmed / repairs requested Boundary: Project Audit system-condition scan, not Change Review of a specific change ### Evidence inspected - ... ### Checklist status | Checklist | Status | Evidence | Confidence | |---|---|---|---| | Tooling/CI drift | ... | ... | ... | | Docs-vs-reality | ... | ... | ... | | Config/env rot | ... | ... | ... | | Dependency health | ... | ... | ... | | Test health/flakiness | ... | ... | ... | | Package/release health | ... | ... | ... | | Instruction/skill drift | ... | ... | ... | Classify findings using the Project Audit priority rubric above: `Critical`, `Watch`, or `Info`. ### Findings 1. Critical / Watch / Info — finding - Evidence: - Impact: - Confidence: High / Medium / Low - Next Keystone module: ### Checks not run - ... ### Overall assessment Health / Watch / At risk / Blocked — rationale ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ``` ## Shared standards For architecture-sensitive or code-quality-sensitive work, load `../../references/engineering-standards.md` and apply it as reference, not dogma.
refactoring4.58 KB
--- name: refactoring description: Program source-code refactoring for an explicit structural code improvement to structure, ownership, types, or duplication that preserves identified executable behavior. Select only when the request requires a structural code change with project-specific invariant proof, or enters through a canonical handoff from another Keystone skill. --- # Refactoring ## Core principle Refactoring changes structure while preserving behavior. Improve design in small, reversible steps, with characterization or regression proof when behavior could drift. ## Load when Load for an existing software project when the user asks to improve code structure while preserving an identified behavior contract: extract or inline code, reduce duplication, clarify names or type boundaries, remove code smells, or reorganize ownership. If the user wants behavior change, use `implementation`. If a failure cause is unknown, use `root-cause-analysis`. If the refactor is broad or risky, create a refactor doc before mutation. At entry, use the full Keystone path when the work changes integrated project code and needs project-specific invariant proof. Handle isolated text cleanup, mechanical formatting, and standalone snippets directly. Explicit invocation selects the full Refactoring behavior. ## Outcome contract A complete refactor reports: - the behavior invariant that must not change; - smells or pressures addressed; - isolation checked before mutation via `../../references/gates/isolation.md`; - characterization/regression proof used before risky edits; - files changed and why each changed; - verification commands/results or explicit proof gaps; - checkpoint handoff to `change-review` when review is needed. ## Process 1. Classify size and risk. - Small/local: one area, clear invariant, existing proof likely enough. - Large/cross-cutting: multiple boundaries, weak coverage, shared contracts, or context-window risk. - Completion criterion: refactor path is either safe for direct mutation or documented first. 2. For large refactors, write a refactor doc under `docs/keystone/refactors/YYYY-MM-DD-<slug>.md` before editing. - Include goal, invariants, smells, affected areas, slices, proof, rollback, and review focus. - Completion criterion: doc is specific enough for `task-creation` or `implementation`. 3. Pass isolation before mutation. - Load/check `../../references/gates/isolation.md`. - Respect unrelated dirty files and protected scope. - Completion criterion: mutation scope is safe. 4. Establish behavior proof. - Prefer existing behavior tests; add characterization coverage when the invariant lacks a tripwire. - Completion criterion: there is a tripwire for accidental behavior change or a documented proof gap. 5. Apply small refactorings. - Prefer rename, extract, inline, move, split, consolidate, simplify conditionals, remove dead code, and clarify ownership. - Keep public contracts stable unless explicitly approved. - Completion criterion: each step is understandable and reversible. 6. Use engineering standards. - Load `../../references/engineering-standards.md` for architecture or ownership decisions. - Remove abstractions that lack current pressure. - Completion criterion: the result has clearer ownership, state, naming, or boundaries. 7. Verify. - Load and pass `../../references/gates/proof.md` for the preserved invariant. - Completion criterion: behavior invariant is supported by observed evidence. 8. Checkpoint and hand off. - Use `../../references/gates/checkpoint.md`. - Hand off to `change-review` for non-trivial refactors or leave an explicit review pointer. ## Smell prompts Investigate duplication, long functions, god objects, feature envy, primitive obsession, data clumps, shotgun surgery, divergent change, hidden control flow, speculative generality, vague managers/helpers, mixed abstraction levels, duplicated state, and unclear ownership. ## Hard rules - Preserve behavior unless the user explicitly approves a behavior change. - Do not disguise feature work as refactoring. - Do not perform broad refactors without a refactor doc. - Do not trust “looks equivalent” without proof or an explicit proof gap. - Do not introduce patterns for imagined futures. ## Output format ```markdown ## Refactor report Invariant: ... Size/risk: small / large Smells addressed: ... Files changed: ... Verification: ... Risks/gaps: ... ### Checkpoint Current skill: refactoring Completed gates: ... Next required skill: change-review / implementation / none Next check: ... Action: continue now / ask user / pending pointer / stop ```
root-cause-analysis13.8 KB
--- name: root-cause-analysis description: Project debugging for a concrete failure in a software or product system. Use for a reproducible or evidence-bearing regression, failing test/build, runtime or integration defect, flake, data corruption, or performance anomaly, or when another Keystone skill needs the cause proven. --- # Root-Cause Analysis ## Core principle Find the root cause before fixing. Root-cause analysis is an evidence ladder: observe the failure, reproduce it, minimize it, trace the mechanism, test falsifiable hypotheses, prove the cause, fix narrowly, guard against regression, verify with exact output, and clean up. No guess-and-check, no cargo-cult edits, no shipping/finalization work. ## Load when Load when the user reports a concrete failure in a software project or product system: a failing test/build, broken runtime behavior, regression, flaky result, performance anomaly, integration failure, suspicious system logs, silent failure, or data corruption. Also load when the task involves: - deciding whether a failure is local, historical, environmental, data-dependent, timing-dependent, or cross-system; - diagnosing nondeterminism, race conditions, retries, timeouts, queues, caches, or distributed boundaries; - interpreting logs/traces/metrics to explain a symptom; - proving whether a suspected fix actually addresses the cause. At entry, use the full Keystone path when the failure belongs to a project artifact or operated product and evidence can be gathered from its code, tests, builds, telemetry, or runtime. Handle general explanations of error messages and non-project troubleshooting directly. Explicit invocation selects the full Root-Cause Analysis behavior. ## Not for - Implementing new behavior unrelated to the failure. - General code improvements without a reproduced problem. - Release finalization, merge strategy, or deployment handoff; use `shipping` after proof and review. - Broad repository audits; use `project-audit`. - Spec decisions where no failure exists; use `product-planning`. - Replacing review, test strategy, or gate validation modules. ## Outcome contract Deliver a root-cause analysis report with: - symptom, impact, affected users/systems, and failure classification; - reproduction steps, exact command/input/environment, or why reproduction was not possible; - minimized failing case when feasible, including what was removed and what still fails; - evidence gathered through logs, tests, code inspection, history, metrics, traces, or instrumentation; - hypotheses considered, which were disproven, and the surviving root-cause hypothesis; - proof that the root cause explains the symptom and predicts observed behavior; - exact fix made or proposed, scoped to the proven cause; - regression test, guard, monitor, or explicit reason none is feasible; - verification commands and exact results/output evidence; - cleanup performed, temporary diagnostics removed, and remaining uncertainty or escalation. Stop only when one of these is true: - the root cause is proven, the narrow fix is verified, and regression protection is in place or justified; - reproduction is impossible after documented attempts and the best available evidence has been preserved; - escalation criteria are met. ## Modes - **Triage:** classify severity, scope, reproducibility, recency, ownership, and risk before fixing. - **Reproduce/minimize:** create the smallest reliable case that demonstrates the failure. - **Instrument/trace:** add temporary logs, probes, traces, assertions, metrics, or diagnostics to observe reality. - **Hypothesis test:** test one falsifiable explanation at a time and record pass/fail evidence. - **Fix:** make the smallest change that addresses the proven root cause. - **Stabilize flaky behavior:** collect repeated runs, isolate nondeterminism, and prove the stabilizing change. - **Performance investigation:** measure baseline, localize bottleneck, prove causality, then optimize narrowly. - **Log/silent-failure investigation:** reconstruct the timeline from logs/traces/state when direct failure output is absent. - **Escalation:** stop local edits and ask for data, access, owner input, incident handling, or risk approval. ## Process 1. **Classify the failure.** Decide whether it is deterministic, flaky, regression, environment-specific, data-dependent, multi-system boundary, performance, silent/log-only, data corruption, security-sensitive, or destructive. Record severity, scope, recency, and owner. 2. **Capture the exact symptom.** Include command, input, error text, expected vs actual, environment, versions, seed/timezone/locale, frequency, affected data, and recent changes. 3. **Reproduce before fixing.** Run the smallest known failing command or scenario. If reproduction is impossible, state the constraint, preserve available evidence, and switch to log/history/data analysis. 4. **Minimize the case.** Reduce inputs, files, flags, mocks, services, data rows, timing windows, browser/device matrix, or integration surface while keeping the failure. Do not minimize away the bug. 5. **Inspect code, data, and history.** Distinguish trigger from cause. If the failure is likely historical, use bisect or targeted history review before guessing. 6. **Trace/instrument narrowly.** Add the smallest temporary diagnostic that can confirm or falsify a hypothesis. Prefer assertions, structured logs, counters, spans, query plans, snapshots, or deterministic seeds over broad logging. 7. **Form falsifiable hypotheses.** A good hypothesis names a mechanism and prediction. Test one at a time. Record disproven hypotheses instead of silently abandoning them. 8. **Prove the root cause.** Show that the cause explains the symptom, reproduces or predicts the failure, and that removing/changing the cause removes the failure. Symptoms alone are not proof. 9. **Fix narrowly.** When this branch will mutate files, load and pass `../../references/gates/isolation.md`, then load `../../references/gates/red.md` and satisfy its observed-red or exception contract before the fix. Change only what the proven cause requires. 10. **Add a regression guard.** Prefer a failing-before/passing-after test. If impractical, add an assertion, monitor, fixture, replay, seed, contract test, migration check, or documented manual proof. 11. **Verify with exact output.** For a mutated fix, load and pass `../../references/gates/proof.md`; run focused verification first, then broader commands if risk warrants. 12. **Clean up.** Remove temporary diagnostics, revert failed experiments, leave useful permanent observability only when justified, and report remaining uncertainty. Branch checklists: - **Cannot reproduce:** verify environment/input parity, collect logs/traces/state, check recent changes, ask for missing data, and escalate if blocked. - **Historical regression:** establish known-good and known-bad revisions with the same command/data/environment; use side-effect-safe bisect only after reproduction is deterministic enough; prove the mechanism in current code before treating the first bad commit as root cause. - **Flaky/race:** run 20/50/100 iterations as risk warrants; vary order, parallelism, clock, seed, async waits, network, cache, filesystem, and shared state; beware instrumentation changing timing. - **Service boundary:** trace correlation/request IDs; compare contracts, schemas, auth, encoding, idempotency, retries, timeout budgets, and partial failures; identify the first system that diverges. - **Performance:** measure baseline and bad metric; localize CPU, memory, IO, network, DB, lock contention, rendering, bundle size, or algorithmic complexity; prove the fix changes that bottleneck. - **Logs/silent failure:** reconstruct a timeline from logs, traces, metrics, state transitions, audit records, and exit codes; check swallowed exceptions, ignored returns, missing awaits, nonzero exits, dropped events, sampling, and log-level/config differences. - **Data corruption:** freeze destructive actions, preserve samples, identify source of truth, writer/read path, migrations/backfills/imports, concurrency, validation, serialization, timezone/locale, precision, blast radius, and rollback safety. Good/bad hypotheses: - **Good:** “The checkout total is doubled because retrying `capturePayment` replays a non-idempotent side effect; if true, two calls with the same request ID will create two ledger rows.” - **Bad:** “Payments are broken.” - **Good:** “The test flakes because it asserts before the debounce timer fires; if true, using fake timers or awaiting the debounce settles it across 100 runs.” - **Bad:** “Probably async weirdness.” - **Good:** “The query slowed after commit X because the new filter prevents index `idx_orders_account_created` from being used; if true, `EXPLAIN` will show a sequential scan and restoring the predicate shape will restore the plan.” - **Bad:** “The database is slow.” Good/bad minimization: - **Good:** Reduce a failing import from a production-sized CSV to three rows that preserve the bad encoding, duplicate key, and null timestamp that trigger the failure. - **Bad:** Replace the import with a mock that no longer exercises parsing, deduplication, or timestamp handling. - **Good:** Reduce a browser failure to one route, one viewport, one user role, and one API response fixture while keeping the visible defect. - **Bad:** Disable authentication, caching, and the API client so the boundary bug disappears. - **Good:** For a flaky test, run the same test alone, in file order, shuffled, with fixed seed, and in parallel to identify the minimal timing/order dependency. - **Bad:** Add arbitrary sleeps until the failure stops appearing once. Escalation/stuck criteria: - Escalate after **three disproven hypotheses** without a stronger next test. - Escalate when there is **no reproduction** and required logs/data/access are unavailable. - Escalate immediately for destructive actions, data-loss risk, security/privacy exposure, production incident impact, legal/compliance concerns, or uncertain repair of corrupted data. - Escalate when the next diagnostic requires credentials, production data, high-cost infrastructure, schema/data mutation, or owner approval. - Escalation output must include symptom, impact, attempts, disproven hypotheses, missing evidence, requested help/access, and safest next action. ## Subagents and reasoning Use subagents for independent root-cause analysis, log review, performance profile interpretation, bisect planning, or hypothesis generation when the active host exposes safe delegation. Use deeper analysis for intermittent, cross-system, security, performance, data-loss, privacy, destructive, or production-impacting failures. When delegation is available, encode required evidence depth and risk standard in the prompt. Subagents may inspect and reason independently, but fixes should converge on one evidence-backed root cause. Ask subagents for competing hypotheses and evidence gaps, not broad code review. When subagents disagree, run the smallest test that distinguishes their explanations. Do not let parallel analysis become parallel guess-and-check edits. ## Hard rules - No fix before evidence supports the root cause. - No “try this” edits unless explicitly labeled as diagnostic experiments and reverted if disproven. - Symptoms are not root causes; keep asking what mechanism produced the symptom. - One hypothesis per experiment; record the prediction and result. - Three disproven hypotheses without progress triggers escalation or a new evidence source. - Regression coverage is required when practical; if impractical, explain why and provide alternate proof. - Temporary instrumentation must be removed or clearly documented before handoff. - Do not declare fixed without command output, exact output proof, metric delta, trace evidence, or equivalent verification evidence. - Do not broaden scope into feature work, refactoring, shipping, finalization, or unrelated cleanup. - For data-loss/security/destructive risk, preserve evidence and escalate before mutation. ## Failure modes - **Guess-and-check spiral:** changing code until symptoms disappear without knowing why. - **Confirmation bias:** keeping the first hypothesis despite contradictory evidence. - **No-op advice:** saying “check logs,” “add tests,” or “investigate further” without specifying which evidence, command, owner, or stop condition. - **Overbroad fix:** refactoring or redesigning more than the failure requires. - **Unreproducible confidence:** claiming success from one weak signal on flaky behavior. - **Minimization that removes the bug:** simplifying the case until the relevant boundary, data shape, or timing condition disappears. - **Instrumentation Heisenbug:** diagnostics change timing, ordering, load, or state enough to hide the failure. - **First-bad-commit tunnel vision:** assuming bisect output is the mechanism without proof. - **Boundary blame ping-pong:** assuming another service owns the issue without request-level evidence. - **Diagnostic litter:** leaving logs, probes, sleeps, flags, generated data, or diagnosis-only state behind. ## Output format ```markdown ## Root-Cause Analysis report Symptom: ... Impact/scope: ... Failure classification: ... Stop condition: fixed / blocked / escalated ... ### Reproduction - Steps/command: ... - Environment/input: ... - Frequency: ... - Minimal case: ... ### Evidence and root cause - Evidence ladder: 1. Observation: ... 2. Trace/minimization: ... 3. Hypothesis test: ... 4. Proof: ... - Disproven hypotheses: ... - Root cause: ... ### Fix - Changed: ... - Why this addresses the cause: ... - Scope intentionally not changed: ... ### Regression protection - Test/guard: ... - Failing-before/passing-after proof or alternate guard: ... ### Verification - Command/result: ... - Exact output proof: ... ### Cleanup / remaining uncertainty - Temporary diagnostics removed: ... - Remaining risk/uncertainty: ... - Escalation needed, if any: ... ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ```
shipping13.1 KB
---
name: shipping
description: Project shipping is authorized immediate delivery execution for identified, already-completed software work. Select only when the user explicitly authorizes an immediate concrete action to commit, push, prepare or create a PR, merge, tag, package, publish, release, deploy, perform an exact repository handoff, or perform destructive cleanup of an exact named completed-delivery target. A canonical handoff qualifies only when it carries the verbatim explicit delivery request and a non-empty exact authorized action set. Context Survey owns pre-authorization cleanup-target reconnaissance while approval is pending.
---
# Shipping
## Core principle
Shipping is deterministic finalization for already-completed work. It proves, reviews, packages, and hands off a release or PR; it never sneaks in new implementation or last-mile fixes.
A shipping decision consumes three shared contracts: load `../../references/gates/proof.md`, `../../references/gates/review.md`, and `../../references/gates/ship.md`. Those files own pass/fail; Shipping gathers evidence and performs explicitly authorized delivery mechanics.
If a final check fails, follow its owning gate's fail action. Final delivery remains blocked until the required gates pass.
## Load when
Load only when the user explicitly asks to deliver completed software project work:
commit it, push it, prepare or create a PR, merge, tag, package, publish, release, deploy, or
perform an exact repository handoff or destructive cleanup of a named completed-delivery
target.
For a Keystone handoff, load `../../references/handoff-packet.md`. The canonical packet's
`evidence` field must carry the verbatim user delivery request and exact authorized
action set. Before acting, derive the effective action set using the packet's
current-intent reconciliation rule. A handoff without carried authorization does not
qualify for Shipping.
At entry, use the full Keystone path when an identified repository change, package,
release candidate, or deployment is complete enough for gates and the user has
explicitly requested a delivery outcome. Handle ordinary summaries, status updates,
and non-project handoffs directly. Explicit skill invocation selects the Shipping
workflow; commit, push, prepare or create a PR, merge, tag, package, publish, release,
deploy, perform an exact repository handoff, or perform destructive cleanup of an
exact completed-delivery target only when that action is in the user's authorized
action set and the gates authorize it.
## Not for
- Starting new implementation or sneaking in last-minute fixes.
- Unfinished code, product, migration, or feature cleanup; route that work to `implementation`. Shipping owns destructive cleanup only for an exact completed-delivery target under explicit user authorization.
- Root-cause debugging; use `root-cause-analysis`.
- Fixing test/build/package failures; use `implementation` for contained repairs or `root-cause-analysis` when the cause is unclear.
- General project risk audits; use `project-audit`.
- Reviewing code quality of a specific change; use `change-review`.
- Shaping unfinished requirements; use `product-planning`.
- Bypassing proof, change-review, package, deploy, or human release approvals.
## Outcome contract
Deliver a strict shipping packet that includes:
- current branch/worktree status and cleanliness;
- scope of completed work and explicit non-scope;
- proof gate evidence with commands/artifacts/results;
- review gate evidence or exact pending review status;
- shipping gate evidence including CI/CD status, deploy preview/staging status when applicable, package/release readiness, rollback plan, and handoff actions;
- multi-target package/release notes where applicable;
- changelog/release-note text or summary;
- unresolved risks and go/no-go verdict;
- authorized action set;
- attempted actions with each exact target and result/status;
- partial completion status, including remaining or unattempted actions;
- recovery guidance for failed or partially completed actions;
- exact human next step.
## Modes
- **PR handoff:** summarize diff, proof, risks, review status, CI status, and reviewer instructions.
- **Release prep:** verify versioning, changelog, build/package artifacts, environment, approvals, rollback, and release command readiness.
- **Deploy handoff:** verify CI/CD pipeline state, deploy preview or staging evidence, environment/feature flag notes, monitoring, and rollback path.
- **Multi-target package/release:** verify each target separately, such as npm/PyPI/GitHub release/Docker/Homebrew/browser extension/mobile artifact, with versions, artifacts, and dry-run evidence.
- **Integration finish:** prepare merge guidance, branch cleanup, post-merge checks, and follow-up owner actions.
- **Delivery packet:** produce final stakeholder notes without changing code.
- **Readiness verdict:** say Shipping / Do not ship / Shipping with risk, backed by gate evidence.
## Process
1. Confirm implementation is complete. If new behavior, fixes, migrations, or unfinished code/product cleanup are still needed, stop and route to `implementation` or `root-cause-analysis`; explicitly authorized destructive cleanup of an exact completed-delivery target remains a Shipping action.
2. Inspect branch/worktree state: branch name, base branch, dirty files, untracked files, commits/diff summary, and whether unrelated changes are present.
3. Define the required gates for this change:
- Load the shared proof, review, and ship gates and identify the evidence each requires for this delivery mode.
- For a handoff, derive the effective authorized action set with the current-intent rule in `../../references/handoff-packet.md`.
4. Run or cite verification evidence. Do not claim passing checks you did not observe. Include command, context, result, and timestamp/context when useful.
5. Check CI/CD awareness: list relevant workflows/pipelines, required checks, latest known status, deploy preview URL or staging environment if available, and any checks not observable locally.
6. Check package/release readiness when applicable: version, changelog, artifact names, package contents, checksums/digests, dry-run output, target registries/platforms, compatibility notes, migration steps, and signing/notarization needs.
7. For multi-target releases, create one evidence row per target. A green web build does not prove a CLI package, Docker image, mobile binary, or plugin package is ready.
8. Confirm rollback and recovery: revert plan, previous version, feature flag/kill switch, database rollback/migration constraints, artifact rollback, owner, and monitoring signals.
9. Prepare PR handoff or release packet: concise summary, scope/non-scope, proof/review/shipping gates, risks, rollout, rollback, and next human actions.
10. Evaluate `../../references/gates/ship.md`.
- On pass, continue to step 11.
- If the review-enablement exception from `../../references/gates/review.md` applies, execute that exception as the terminal delivery-action branch, skip step 11, and continue directly to step 12. After the checkpoint, stop with the pending Review Gate handoff.
- For any other failed prerequisite, follow its owning gate's fail action, include the failed evidence, and route to the correct module.
11. After every prerequisite gate passes, perform only actions in the user-authorized action set.
- Before each action, reconcile the effective action set with the latest explicit user instructions via `../../references/handoff-packet.md`.
- Before a workspace mutation, load and pass `../../references/gates/isolation.md`. For destructive cleanup, pass its Requested cleanup mode.
- Resolve the exact targets before acting: repository, branch, remote, PR, tag, package, registry, release, environment, artifact, handoff recipient, or cleanup path as applicable.
- For destructive cleanup, require target confirmation: re-confirm the exact paths, branches, or artifacts, the explicit request, and recoverability or rollback; prefer recoverable operations.
- When a package action creates or changes an artifact, inspect the actual artifact contents, checksum, signature, and target-specific checks. Re-run the applicable proof and ship gates before any dependent publish or release action; proceed only when both gates pass on that artifact.
- Record each action and its result.
- Stop when any action fails, report partial completion, and leave unauthorized or unattempted remaining actions untouched.
12. Run the checkpoint gate and end with a clear verdict: Shipping, Do not ship, or Shipping with risk. If any gate is missing, the checkpoint action is not `stop`; route or prompt for the next required module/check.
## Subagents and reasoning
Use subagents for bounded release-note drafting, checklist verification, artifact inspection, CI/CD status inspection, package manifest review, or independent review of the shipping packet when the active host exposes safe delegation. Use deeper analysis for multi-platform packaging, production releases, security-sensitive changes, migrations, deploys with customer impact, or unresolved release risk. When delegation is available, encode required evidence depth and release standard in the prompt. Subagents must not introduce new implementation.
## Hard rules
- No new implementation in shipping mode. Gate failures create a handoff, not stealth fixes.
- Enforce proof, change-review, and shipping gates. Do not collapse them into one vague readiness statement.
- Evidence before assertions: every readiness claim needs command output, artifact proof, CI/CD status, deploy preview/staging proof, or documented review.
- Do not bypass review or package gates because the change “looks small.”
- Keep changelog/release notes user- or operator-relevant; avoid dumping raw commit noise.
- State branch name and working tree cleanliness when available.
- If readiness is uncertain, say “Do not ship” or “Shipping with risk,” not “done.”
- Never commit, push, prepare or create a PR, merge, tag, package, publish, release, deploy, perform an exact repository handoff, or perform destructive cleanup unless the user explicitly requested that action and the gates support it.
- Run the checkpoint gate before the final response; a shipping packet with missing gates must name the next event instead of sounding final.
## Failure modes
- **Victory lap without proof:** announcing completion before tests/build/review evidence.
- **Last-mile coding:** making new fixes under the cover of release prep.
- **Stealth release:** tagging, publishing, deploying, or merging without explicit approval.
- **CI blindness:** relying only on local checks while required CI/CD, preview, or staging is red or unknown.
- **Single-target tunnel vision:** treating one package/build target as proof for all targets.
- **Rollback omission:** shipping without a practical revert, rollback, or recovery path.
- **Release-note mush:** vague notes that omit impact, migration, rollout, or risk.
- **Dirty handoff:** leaving untracked files, unclear branch state, or hidden manual steps.
- **Gate theater:** listing checks without results or timestamps/context.
- **Terminal ambiguity:** ending with no clear human action, rollback owner, or next Keystone module when shipping is not cleanly complete.
## Output format
```markdown
# Shipping packet
Verdict: Shipping / Do not ship / Shipping with risk
Branch/status: ...
Gate summary: Proof ... / Change Review ... / Shipping ...
### Authorized delivery actions
- Authorized action set: ...
- Attempted actions, exact targets, and result/status:
| Action | Exact target | Result/status |
|---|---|---|
| ... | ... | succeeded / failed / not attempted |
- Partial completion: none / details
- Remaining actions / unattempted actions: ...
- Recovery guidance: ...
- Next step: ...
### Scope
- Completed: ...
- Not included: ...
### Proof gate
| Evidence | Good/Bad | Result | Notes |
|---|---|---|---|
| Good: `npm test` observed exit 0 on this branch | Good | Pass | Include command/output summary |
| Bad: “tests should pass” without running or CI link | Bad | Missing | Not acceptable evidence |
### Change Review gate
- Good evidence: approved PR review, completed self-review checklist, security/design approval when required.
- Bad evidence: “looks fine,” assumed approval, or stale review from before major changes.
- Status: ...
### Shipping gate
- CI/CD: workflow/status/link or not observable and why.
- Deploy preview/staging: URL/environment/check result or not applicable.
- Package/release targets: versions, artifacts, checksums/digests, dry runs, compatibility notes.
- Rollback plan: exact revert/redeploy/unpublish/feature-flag path and owner.
### Release notes / changelog
- ...
### Risks and aborts
- Failed/missing gates:
- Risks accepted:
- Route if not shippable: implementation / root-cause-analysis / change-review / project-audit
### PR handoff / release packet
- Summary:
- Proof:
- Change Review:
- Rollout:
- Rollback:
- Next human actions:
### Checkpoint
Use the required fields from `../../references/gates/checkpoint.md`.
```
## Explicit-only finalization
Commit, push, PR, merge, tag, package, publish, release, deploy, repository handoff, and destructive cleanup require explicit user request. Preparing notes is allowed; performing the action is not implicit.
task-creation14.7 KB
--- name: task-creation description: Project delivery breakdown for a concrete software or product goal. Use when that goal needs implementation slices, milestones, dependencies, verification gates, or agent-ready work, or when another Keystone skill hands off a shaped goal for sequencing. --- # Task Creation ## Core principle Task Creation turns a goal into sequenced, reviewable vertical slices of work. A good task-creation makes implementation easier because every slice has a visible result, clear constraints, and a verification path. It makes review easier because reviewers can compare changes against stated goals, requirements, risks, and acceptance checks. Task Creation is sequencing, not execution. A task-creation is not proof; only inspected changes, tests, demos, or other evidence prove completion. ## Load when Load for a concrete software or product project when the user asks to: - break down a feature, fix, refactor, migration, tool, system, or project - produce milestones, implementation steps, tickets, issues, vertical slices, or phases - decide sequencing, dependencies, iterations, scope cuts, or parallelization - turn an approved or sufficiently shaped goal into implementable work - prepare work for coding agents, reviewers, or subagents - sequence greenfield architecture after the core users, runtime, and tradeoffs are stable enough to slice - split a large task into reviewable chunks without exposing another public command Also load when implementation is requested but the goal is broad enough that coding immediately would hide major sequencing decisions. If behavior, scope, UX, or architecture tradeoffs are still undecided, route to `product-planning` first; use Task Creation once the desired outcome is stable enough to sequence. At entry, use the full Keystone path for project delivery work tied to a repository, product initiative, approved specification, or engineering lifecycle outcome. Handle standalone lists, personal planning, and ordinary to-do decomposition directly. Explicit invocation selects the full Task Creation behavior. ## Not for Do not use Task Creation for: - implementation, file edits, refactors, migrations, or generated code - debugging a known failure; use `root-cause-analysis` first, then return to Task Creation if a repair sequence is needed - context-survey-only tasks; use `context-survey` first when facts are missing - copy/design shaping as the primary work; use `product-planning` first when output is prose, UX, or design direction - final verification, release readiness, or completion claims; use `change-review`, `project-audit`, or `shipping` ## Outcome contract A Task Creation output must include: 1. Goal: the intended outcome in one or two sentences. 2. Context inspected: files, docs, requirements, or assumptions used. 3. Requirements inventory: critical requirements, non-functional requirements, good-to-haves, constraints, and open questions. 4. Stack and architecture context: current or proposed technologies, boundaries, integrations, and limitations. 5. Iteration layering: iteration 1, iteration 2, iteration 3, etc., with explicit scope cuts. 6. Vertical slices: ordered work items that each deliver an end-to-end outcome. 7. Verification gates: how each slice can be tested, reviewed, or demonstrated. 8. Risks and dependencies: what can block, invalidate, or reorder the work. 9. Handoff: recommended next primary module and any subagent/analysis-depth suggestions. If information is missing, state the assumption or ask the smallest set of questions required to avoid a bad task-creation. ## Modes ### Feature task-creation Use for new capabilities in an existing product. Focus on the user/operator goal, existing entry points, data paths, APIs, UI surfaces, tests, deployment constraints, and the smallest end-to-end slice that proves the feature path. Add progressive enrichment after the core loop works. Avoid horizontal buckets like "database", "backend", "frontend", and "tests" unless they are nested inside a vertical slice. ### Greenfield architecture task-creation Use for new projects, tools, services, apps, packages, or substantial standalone systems. Start architecture-first: name the primary users, runtime, language, framework, hosting, storage, auth, observability, CI, packaging constraints, system boundaries, and integration points. Iteration 1 should prove the architecture can run, test, deploy, and support one meaningful vertical path, not merely create folders. ### Refactor/migration task-creation Use when behavior should remain stable while internals change. Focus on the current behavior contract, compatibility expectations, affected surfaces, consumers, migration seams, adapters, flags, dual-run paths, rollback strategy, observability, and characterization tests. Prefer strangler-style or seam-first slices over broad rewrites. ### Subagent-parallel task-creation Use when independent workstreams can proceed safely in isolated workspaces. Map the dependency graph before delegation. Name shared files, interfaces, merge-risk hotspots, delegation purpose, required analysis depth, context packets, expected artifacts, integration order, and review checkpoints. Only parallelize slices that can be verified independently or integrated behind a clear contract. ## Process 1. Identify the goal. - Restate the desired end state, not just the requested activity. - Identify who benefits and what observable change proves value. - If the goal is unclear and cannot be inferred from context, ask one focused question. 2. Inspect before asking broad questions. - Read relevant files, docs, issues, architecture notes, tests, package manifests, and existing modules when available. - Use `context-survey` for unfamiliar libraries, external constraints, or repository-wide discovery. - Ask clarifying questions only when inspection cannot resolve a decision that materially changes the task-creation. 3. Build the requirements inventory. - Separate critical requirements from non-functional requirements and good-to-haves. - Capture explicit user constraints, protected files, deadlines, compatibility needs, and review expectations. - Mark assumptions and unknowns instead of silently inventing requirements. 4. Identify stack and constraints. - Note languages, frameworks, package managers, deployment targets, test tooling, persistence, auth, APIs, and platform limits. - For greenfield work, propose architecture choices only after naming tradeoffs and constraints. - For existing systems, prefer the established stack unless the goal requires a change. 5. Choose iteration layers. - Define iteration 1 as the smallest coherent outcome that proves the path. - Define later iterations as progressively richer outcomes, not random leftovers. - Make explicit what is intentionally deferred. 6. Slice vertically. - Each slice should cross needed layers to produce a reviewable result. - Include data/model/API/UI/test/docs work inside the slice when needed for that result. - Avoid phases that finish entire subsystems before any end-to-end value appears. 7. Add verification gates. - Give each slice an acceptance check, test command, manual review path, or demo criterion. - Include regression checks for migrations and refactors. - State when review should happen and what reviewers should inspect. 8. Prepare the handoff. - Recommend the next Keystone module. - Identify subagent opportunities, required context, and analysis depth. - Call out risks, dependencies, and open questions that should block implementation if unresolved. ## Requirements inventory Use this inventory before sequencing work: - Goal: what outcome the user wants. - Critical requirements: must be true or the work fails. - Non-functional requirements: performance, security, reliability, accessibility, privacy, maintainability, observability, compatibility, cost, release, and operational concerns. - Good-to-haves: valuable but deferrable enhancements. - Constraints: protected files, APIs, tech stack, deadlines, repo conventions, deployment targets, team/process limits. - Current state: what exists now, with file/source references when available. - Unknowns: decisions or facts still unresolved. - Assumptions: temporary beliefs used to proceed. When requirements conflict, surface the conflict before writing slices. ## Iteration layering Iteration layers should describe increasing confidence and capability: - Iteration 1: skeleton or core path. One thin, end-to-end result that validates the architecture, integration point, or user journey. - Iteration 2: completeness and resilience. Add common cases, validation, error states, compatibility, and tests around the proven path. - Iteration 3: polish and scale. Add edge cases, performance, accessibility, observability, docs, migration cleanup, and nice-to-haves. - Later iterations: optional expansion, hardening, automation, or product refinements. For greenfield projects, iteration 1 must include a runnable or executable foundation plus one meaningful vertical path. For migrations, iteration 1 should establish safety: characterization checks, seams, adapters, or observability before broad movement. ## Task quality bar Each task or slice must have: - a name that describes the delivered outcome - user/operator/developer value - inputs and dependencies - files or areas likely involved, when known - exact acceptance criteria - verification method - rollback or safety note when risk is non-trivial - review focus: what a reviewer should inspect A weak task says "implementation backend" or "add tests". A strong slice says "Persist saved searches end-to-end behind the existing search UI, with API validation, storage migration, and regression coverage for loading saved searches." ## Subagents and reasoning Use subagents when sequencing benefits from independent context gathering, critique, or parallel workstream design and the active host exposes safe delegation: - read-only context-survey on separate code areas, external APIs, or prior art - architecture critique for greenfield foundations and migrations - risk critique for security, data loss, compatibility, or release sequencing - implementation delegation only after slices are independent and interfaces are stable For each proposed delegation, specify purpose, required analysis depth, context packet, expected output artifact, files or areas off limits, and integration/review checkpoint. When delegation is available, encode required evidence depth and risk standard in the prompt. Do not use subagents to bypass ambiguity. Resolve shared interfaces and sequencing first. ## Hard rules - Task Creation stops before implementation. - Task Creation mutates only the task artifact selected by the Task artifact rule. - Do not rename `task-creation` to `plan`. - Do not expose `/plan`. - Do not claim the task-creation proves completion. - Do not skip goal identification. - Do not ask broad clarifying questions before inspecting available context. - Do not produce horizontal-only task-creations. - Do not hide assumptions, unresolved questions, or conflicts. - Route risky task sets through `implementation` with explicit verification gates and change-review points. ## Failure modes - Activity list: names actions but never states the outcome. - Horizontal buckets: separates backend/frontend/tests so no slice is independently valuable. - Big-bang architecture: designs everything before proving one runnable path. - Faux certainty: treats assumptions as facts. - Question spam: asks what inspection could answer. - Implementation leak: starts coding, editing files, or choosing exact code structure beyond sequencing needs. - Change Review-hostile output: lacks acceptance criteria, test commands, or reviewer focus. - Parallelism theater: delegates coupled workstreams that collide on shared files or undefined interfaces. - Task Creation-as-proof: reports success because a task-creation exists. ## Output format Use this structure unless the user requested a different artifact. For small contained tasks, compress sections while preserving goal, assumptions, slices, verification, risks, and handoff: ```markdown # Task Creation: <goal> ## Goal <one or two sentences describing the desired outcome> ## Context inspected - <files, docs, issues, sources, or "none available"> ## Requirements inventory ### Critical requirements - <must-have> ### Non-functional requirements - <quality/operational constraint> ### Good-to-haves - <deferrable enhancement> ### Stack and architecture context - Stack: <technologies, frameworks, platforms, deployment targets> - Boundaries: <modules, layers, services, UI/data/domain seams> - Integrations: <APIs, persistence, queues, auth, payment, external systems> - Limitations and constraints: <protected files, compatibility, process, deadline, cost, operational limits> ### Unknowns and assumptions - Unknown: <question that matters> - Assumption: <assumption used for this task-creation> ## Iteration layering ### Iteration 1: <core path / architecture skeleton / safety seam> - Outcome: <reviewable result> - Scope: <included> - Deferred: <not included> ### Iteration 2: <complete common cases> - Outcome: <reviewable result> - Scope: <included> - Deferred: <not included> ### Iteration 3: <hardening / polish / scale> - Outcome: <reviewable result> - Scope: <included> - Deferred: <not included> ## Vertical slices 1. <slice name> - Value: <who benefits and how> - Work: <end-to-end changes at a sequencing level> - Dependencies: <prior slices or decisions> - Acceptance: <observable completion criteria> - Verification: <test/review/demo method> - Change Review focus: <what reviewers should inspect> ## Risks and dependencies - <risk, impact, mitigation> ## Subagent opportunities - <delegation purpose, required analysis depth, context packet, expected artifact> ## Handoff Next module: `<context-survey|implementation|root-cause-analysis|change-review|project-audit|shipping|product-planning>` because <reason>. ### Checkpoint Use the required fields from `../../references/gates/checkpoint.md`. ``` Keep the output detailed enough to guide implementation and review, but short enough that each slice remains actionable. ## Task artifact rule Slice the work naturally before counting top-level vertical slices. Then choose delivery: - For 1–5 top-level vertical slices, return the full task creation in conversation without creating a file. - For 6 or more top-level vertical slices, write the task creation under `docs/keystone/tasks/YYYY-MM-DD-<slug>.md`. - An explicit request for chat only or conversation only keeps any length in conversation. - An explicit request to write or save the task creation produces the artifact at any length. Explicit delivery requests override the numeric threshold. When writing an artifact, return a concise summary and its path in conversation; do not duplicate the full task creation there.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- MIT
- Package author
- static-var
- Keywords
- See publisher keywords
Declared capabilities
- Read
- Write
- Review
- Workflow
Package observed Oct 3, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 3, 2026 · 18:00 UTC
- Collection status
- Collected
plugins_6a512fae91f881918208460aae465ff9
Download plugin data (JSON)Before you connect Keystone
How do I connect it?
Open the publisher's marketplace listing to check current availability and follow its connection instructions. This directory does not install plugins. Check the requested access and any account requirements before connecting.
Check marketplace availability ↗
Does it require paid access?
We have not established the pricing or subscription requirements for this plugin. An absent price does not mean free access.
Compare researched pricing and access models →
How can I evaluate it?
Check the declared skills and available files, then try a small task whose result you can verify. Our archived descriptions and instructions establish publisher claims, not tested runtime quality. Review sources and coverage limits.