{"id":16448,"plugin_id":"plugins_6a512fae91f881918208460aae465ff9","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:13:26.062Z","digest":"eef9008a0a8d67bb708b2c874c9ba968a2514b1844ed30e3d8736cbdb23fc981","against":null,"payload":{"description":"Project change review for a concrete diff, branch, PR, patch, migration, fix, implementation result, or project plan that needs a new evidence-backed readiness verdict, or when another Keystone skill needs the Review Gate evaluated.","included_files":[],"name":"change-review","skill_md_contents":"---\nname: change-review\ndescription: Project change review for a concrete diff, branch, PR, patch, migration, fix, implementation result, or project plan that needs a new evidence-backed readiness verdict, or when another Keystone skill needs the Review Gate evaluated.\n---\n\n# Change Review\n\n## Core principle\nChange Review is an independent, read-only attempt to disprove readiness.\n\nAsk two questions at the same time:\n1. **Spec axis:** does the work satisfy the stated requirements and acceptance criteria?\n2. **Standards axis:** is it secure, correct, maintainable, tested, and safe to operate?\n\nDo not assume changed lines are the blast radius. Trace callers, callees, contracts,\ndata flow, tests, runtime paths, and user impact before giving a verdict.\n\n## Load when\nLoad when a concrete software project artifact needs review: a diff, branch, PR, patch,\nmigration, fix, implementation result, or project plan that must be assessed against\nits specification and engineering standards.\n\nAlso load when another Keystone skill needs `../../references/gates/review.md` satisfied before shipping.\n\nAt entry, use the full Keystone path when there is an inspectable project artifact and\na readiness or blocker verdict is the outcome. Handle general critiques, prose review,\nand standalone explanations directly. Explicit invocation selects the full Change Review behavior.\n\n## Not for\nDo not use Change Review for:\n- fixing, refactoring, formatting, or rewriting code\n- committing, merging, tagging, publishing, or shipping\n- initial implementation planning before a reviewable artifact exists\n- open-ended context-survey with no concrete artifact to assess\n- debugging where the requested outcome is a fix\n\nIf asked to review and fix, review first, stop, and hand findings to `implementation`, `root-cause-analysis`,\n`context-survey`, `shipping`, or a human only after explicit permission.\n\n## Outcome contract\nA complete review returns:\n- verdict: **Block**, **Caution**, or **Looks good**\n- findings ordered P0, P1, P2, P3, then Nitpicks\n- evidence for every finding: file/line, behavior path, contract, test, log, or doc\n- user impact and why the severity is justified\n- remediation guidance without applying the fix\n- tests that should be added or updated for affected behavior\n- scope reviewed, validation run, limitations, and read-only confirmation\n\nThe review is incomplete if it only inspects the diff, only comments on style, or\ncannot explain how the work behaves at runtime.\n\n## Change Review passes\nPerform multiple passes. New evidence from one pass expands later passes.\n### Pass 0: scope and baseline\n- Identify artifact reviewed: diff, branch, files, release candidate, or plan result.\n- Read the user request, issue, spec, acceptance criteria, and claimed completion.\n- Check repository status without modifying files.\n- Record uncommitted work as context, not cleanup.\n\n### Pass 1: spec compliance\n- Compare implementation against explicit requirements and non-goals.\n- Check edge cases, error states, and acceptance criteria.\n- Separate spec misses from standards concerns.\n- Treat a clean implementation of the wrong behavior as a finding.\n\n### Pass 2: correctness and runtime paths\n- Trace primary success and failure paths end to end.\n- Follow changed functions into helpers, services, adapters, persistence, UI, jobs, and\n  serializers.\n- Validate inputs, outputs, invariants, state transitions, retries, ordering,\n  concurrency assumptions, and error propagation.\n- Look for nullability, off-by-one, time, encoding, pagination, caching, idempotency,\n  cancellation, and partial-failure issues.\n\n### Pass 3: regression and compatibility\n- Identify callers, consumers, and workflows that rely on old behavior.\n- Check public APIs, CLIs, schemas, migrations, persisted data, environment variables,\n  feature flags, configuration defaults, and documentation.\n- Consider rollback, downgrade, mixed-version, and incremental rollout risks.\n- Search for tests or fixtures that encode previous behavior.\n\n### Pass 4: security, privacy, and abuse resistance\n- Review authentication, authorization, tenancy, secrets, logging, validation,\n  injection, XSS, SSRF, path traversal, unsafe deserialization, and RCE surfaces.\n- Check whether sensitive data leaks through errors, logs, telemetry, URLs, caches,\n  exports, screenshots, or third-party calls.\n- Consider malicious users, compromised clients, replay, races, resource exhaustion,\n  privilege escalation, and denial of service.\n\n### Pass 5: tests and proof\n- Map changed behavior to existing tests.\n- Identify missing unit, integration, contract, regression, migration, security,\n  accessibility, performance, or end-to-end coverage.\n- Prefer behavior assertions over implementation trivia.\n- Run focused read-only validation when practical: existing tests, type checks, lint,\n  builds, or targeted commands.\n- If validation cannot run, state why and what should be run.\n\n### Pass 6: maintainability and architecture\n- Assess clarity, cohesion, naming, dependency direction, duplication, complexity,\n  observability, and debuggability.\n- Check architectural boundaries, local conventions, and API contracts.\n- Flag brittle abstractions, hidden coupling, unnecessary cleverness, and premature\n  generalization when they create real maintenance risk.\n\n### Pass 7: user impact and final consistency\n- Translate technical issues into affected personas, workflows, data, accessibility,\n  performance, reliability, and support burden.\n- Re-rank findings by blast radius, likelihood, recoverability, and detectability.\n- De-duplicate findings, verify evidence, and state limitations honestly.\n\n## Severity rubric\nSeverity reflects realistic impact, not fix size.\n\n### P0: Critical blocker\nImmediate or likely severe harm. Examples:\n- data loss, corruption, or irreversible destructive action\n- unauthorized access, privilege escalation, secret exposure, or major privacy breach\n- production outage or release artifact that cannot safely deploy\n- legal/compliance risk with material impact\n\nP0 means do not ship or merge without accountable human acceptance and mitigation.\n\n### P1: Blocking defect\nHigh-impact issue that violates core requirements or creates serious regression risk.\nExamples:\n- primary workflow broken for a meaningful user segment\n- incorrect billing, permissions, persistence, or business logic\n- migration or compatibility gap that can break real deployments\n- high-risk behavior lacking tests plus a plausible failure mode\n\nP1 normally blocks shipping.\n\n### P2: Important non-blocker or conditional blocker\nMaterial issue with bounded impact, lower likelihood, or workaround. Examples:\n- edge case with clear user impact\n- moderate-risk test gap\n- maintainability issue likely to cause near-term bugs\n- weak observability for a risky path\n\nState whether release context makes it blocking.\n\n### P3: Low-risk improvement\nValid concern with limited impact. Examples:\n- confusing name or local complexity that slows future work\n- minor non-hot-path performance inefficiency\n- incomplete docs for non-critical behavior\n- small test organization weakness\n\nP3 should not block unless it compounds with related risks.\n\n### Nitpick\nCosmetic, preference-level, or optional feedback: unenforced formatting, wording tweaks,\nor style suggestions with no correctness or maintainability impact. Keep nitpicks\nseparate from severity findings.\n\n## Impact tracing\nFor each meaningful change, trace:\n- **Entry points:** user action, API route, CLI, job, event, hook, or import.\n- **Callers:** who invokes this and what assumptions they make.\n- **Callees:** helpers, libraries, persistence, network calls, and side effects.\n- **Data flow:** input, validation, transformation, storage, serialization, output.\n- **Contracts:** types, schemas, public APIs, flags, config, docs, and errors.\n- **Runtime paths:** success, failure, retry, timeout, cancellation, concurrency.\n- **Tests:** existing coverage, missing assertions, fixtures, mocks, snapshots.\n- **Users:** visible behavior, accessibility, performance, reliability, trust.\n\nIf tracing leaves uncertainty, gather more read-only evidence or report the limitation.\nDo not invent confidence.\n\n## Security and regression checklist\nAsk for every non-trivial review:\n- Can a user access, modify, infer, or delete data they should not?\n- Are authn, authz, tenancy, and ownership checked at the right layer?\n- Can untrusted input reach queries, interpreters, shells, paths, templates, redirects,\n  or deserializers unsafely?\n- Are secrets, tokens, PII, or internal identifiers exposed in logs, errors, telemetry,\n  URLs, caches, or client bundles?\n- Did defaults, permissions, feature flags, or safeguards become unsafe?\n- Are races, duplicate submissions, retries, replay, and out-of-order events safe?\n- Can persisted data be corrupted, stranded, or made hard to rollback?\n- Are public APIs, stored data, configs, and integrations backward compatible?\n- Does failure degrade safely without hidden partial success?\n- Are performance, resource use, accessibility, localization, and platform differences\n  acceptable for realistic users and abuse?\n- Do tests cover the affected behavior and important regression paths?\n\n## Subagents and reasoning\nUse read-only subagents for separable risks: security/privacy, test coverage,\narchitecture/API compatibility, persistence/migration, accessibility/user impact,\nperformance, concurrency, or release risk. Use deeper analysis for security-sensitive,\ndata-loss, billing, permissions, public API, migration, or cross-system reviews. When delegation is available, encode required evidence depth and review standard in the prompt.\n\nSubagents must receive the read-only contract and return evidence-backed findings, not\npatches. Reconcile duplicates and conflicts before reporting. The primary reviewer\nowns final severity and verdict.\n\n## Hard rules\n- Read-only: inspect files without editing, formatting, generating, staging, committing, merging, tagging, or publishing them.\n- Report issues without silently fixing them.\n- Do not run destructive or project-mutating commands.\n- Do not rely only on changed lines; inspect impacted code paths and contracts.\n- Do not approve solely because tests pass.\n- Do not report speculation as fact; mark uncertainty.\n- Do not bury blockers under minor comments.\n- Do not disguise style preferences as correctness findings.\n- Do not omit needed tests when behavior changed.\n- Do not satisfy `../../references/gates/review.md` unless blockers and non-blockers are separated.\n- Load `../../references/gates/review.md` before the verdict; Change Review owns the review execution and supplies the gate evidence.\n- Run the checkpoint gate before the final response; if review passes and delivery/finalization was requested, route to `shipping` or leave an explicit shipping prompt.\n\n## Failure modes\nAvoid these anti-patterns:\n- **Single-pass skim:** one read of changed lines plus generic comments.\n- **Diff tunnel vision:** missing callers, callees, contracts, and user impact.\n- **Checklist theater:** naming security/tests without tracing actual risk.\n- **Green-test rubber stamp:** assuming current tests prove new behavior.\n- **Spec blindness:** judging code quality while requirements are unmet.\n- **Standards blindness:** accepting unsafe or fragile code because the narrow spec passes.\n- **Severity inflation:** turning preferences into blockers.\n- **Severity deflation:** downgrading real user harm because the fix is small.\n- **Patch creep:** fixing, refactoring, or committing instead of reviewing.\n- **Unowned uncertainty:** failing to state what was not verified.\n- **Lost next event:** Change Review passes but never routes or prompts for `shipping` when finalization remains.\n\n## Output format\nWorked finding example:\n```markdown\n### P1\n- Missing tenant check on invoice export\n  - Evidence: `api/exportInvoice.ts:42` accepts `invoiceId` and loads the invoice without comparing `invoice.accountId` to the authenticated account; `/invoices/:id/export` is reachable by any logged-in user.\n  - Impact: A user who guesses another invoice ID can download billing data from a different account, which is a privacy and authorization breach.\n  - Recommendation: Enforce tenant ownership before export and return the existing unauthorized response on mismatch.\n  - Tests needed: Add an integration test where account A requests account B's invoice and receives 403/no file, plus a happy-path same-account export test.\n```\n\nUse this structure:\n```markdown\n## Verdict\nBlock | Caution | Looks good\n\n## Scope reviewed\n- Artifact reviewed:\n- Key files/paths inspected:\n- Validation run:\n- Review limitations:\n- Read-only confirmation: no files changed by this review\n\n## Findings\n### P0\n- [Title]\n  - Evidence:\n  - Impact:\n  - Recommendation:\n  - Tests needed:\n### P1\nNone\n\n### P2\nNone\n\n### P3\nNone\n\n## Nitpicks\nNone\n\n## Tests to add or update\n- Behavior:\n  - Suggested coverage:\n  - Why it matters:\n## Handoff\n- Blockers:\n- Non-blocking follow-up:\n- Suggested owner module: implementation, root-cause-analysis, context-survey, shipping, or human\n\n### Checkpoint\nUse the required fields from `../../references/gates/checkpoint.md`.\n\n```\n\nIf a severity has no findings, write `None`. Recommendations must be actionable but\nmust not be applied by Change Review.\n\n## Shared standards\n\nFor architecture-sensitive or code-quality-sensitive work, load `../../references/engineering-standards.md` and apply it as reference, not dogma.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}