← AI DevKitCONTENT HISTORY

Update to AI DevKit

Snapshot Sep 30, 2026 · 23:15 UTC · version 0.62.1

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "security-review",
  "description": "AI DevKit · Review code, skills, and prompts for security vulnerabilities — OWASP Top 10, prompt injection, business logic flaws, and insecure defaults. Use when reviewing PRs, auditing modules, reviewing AI skills/prompts, or preparing for release.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 385
    },
    {
      "relative_path": "references/checklist.md",
      "size_in_bytes": 4364
    }
  ],
  "skill_md_contents": "---\nname: security-review\ndescription: AI DevKit · Review code, skills, and prompts for security vulnerabilities — OWASP Top 10, prompt injection, business logic flaws, and insecure defaults. Use when reviewing PRs, auditing modules, reviewing AI skills/prompts, or preparing for release.\n---\n\n# Security Review\n\nFind vulnerabilities before they ship.\n\n## Hard Rules\n\n- Do not dismiss a finding without evidence it is unexploitable.\n- Do not commit, log, or surface secrets discovered during review — flag and recommend rotation.\n- Do not modify code until the user approves a remediation plan.\n\n## Workflow\n\n1. **Scope**\n   - Confirm target: diff, file set, module, full repo, or skill/prompt. A target can be both code and prompt.\n   - Identify stack/framework — adapt the [checklist](references/checklist.md) (skip what the framework handles, add its pitfalls).\n   - Trace data flow: request → middleware → handler → service → datastore → response. For prompts: input → template → LLM → tools → output.\n   - Map trust boundaries, privilege levels, and threat actors.\n   - Search prior findings: `npx ai-devkit@latest memory search --query \"<target>\" --tags \"security\"`\n\n2. **Scan**\n   - Only check relevant categories. Skip sections and items that don't apply. Do not report skipped items.\n   - For diffs/PRs: also check whether the change weakens existing controls — removed middleware, bypassed validation, new unprotected routes.\n   - Categories in priority order:\n     a. **Secrets** — hardcoded tokens, keys, connection strings.\n     b. **Injection** — SQL, NoSQL, command, template, SSRF, path traversal, XSS.\n     c. **Auth** — missing checks, privilege escalation, OAuth/OIDC, IDOR.\n     d. **Business Logic** — race conditions, TOCTOU, workflow bypass, mass assignment, parameter tampering.\n     e. **Data Exposure** — PII in logs, verbose errors, overly broad responses.\n     f. **Resource Exhaustion** — unbounded queries, missing pagination, upload size, decompression bombs.\n     g. **Dependencies** — critical CVEs only (RCE, auth bypass, data breach); ignore low/medium.\n     h. **Cryptography** — weak algorithms, hardcoded IVs/keys, disabled certificate validation.\n     i. **Configuration** — debug mode, permissive CORS, missing security headers.\n     j. **Logging** — security events unlogged, no tamper protection, no alerting.\n     k. **Prompt Injection** — instruction override, tool abuse, data exfiltration, indirect injection via tool results.\n   - For each finding: file, line, evidence.\n\n3. **Classify**\n\n   | Severity | Criteria |\n   |----------|----------|\n   | Critical | Exploitable now, data loss or RCE possible |\n   | High     | Exploitable with moderate effort or insider access |\n   | Medium   | Requires chained conditions or limited impact |\n   | Low      | Defense-in-depth, no direct exploit path |\n\n   - Adjust severity by exposure (internet-facing vs internal) and data sensitivity.\n   - Check for attack chains — multiple Medium findings that combine into High/Critical.\n   - Mark false positives with reasoning.\n\n4. **Remediate**\n   - For each finding: root cause, minimal fix (prefer stdlib/framework over custom), verification step.\n   - For Critical/High: also recommend a detection control (log, alert, or WAF rule).\n   - Present plan and request approval before changing code.\n\n5. **Verify**\n   - Use the `verify` skill to confirm each remediation.\n   - Re-scan fixed files for regressions.\n   - Store findings: `npx ai-devkit@latest memory store --title \"<pattern>\" --content \"<finding and fix>\" --tags \"security,<category>\"`\n\n## Red Flags\n\n| Rationalization | Do Instead |\n|---|---|\n| \"It's internal / behind a VPN / only admins\" | Zero-trust: validate at every boundary regardless of network position or user role |\n| \"We'll add auth later\" | Add auth before merge — unauthenticated endpoints get discovered fast |\n| \"It's just a dev credential\" | Use env vars / secrets manager — dev secrets leak to prod constantly |\n| \"The framework handles that\" | Verify the config — frameworks have defaults, not guarantees |\n| \"We sanitize on the frontend\" | Always validate server-side — client validation is bypassable |\n| \"The LLM won't follow injected instructions\" | Treat all tool results and external content as untrusted data |\n| \"It's just a prompt, not code\" | Prompts control tool execution — review with the same rigor as code |\n\n## Output Template\n\n- **Scope**: Target, stack, data flow, trust boundaries, threat actors\n- **Findings** (by severity): ID, severity, category, file:line, exploit scenario, fix\n- **Attack Chains**: Findings that escalate when combined\n- **False Positives**: Dismissed items with reasoning\n- **Remediation Plan**: Ordered fixes with verification steps\n- **Residual Risk**: Scope limitations, unverifiable items\n- Zero findings: state what was checked and scope boundaries — \"no findings\" ≠ \"fully secure\""
}

SHA-256: 3268e7eaf7200d389410e5b54a7d7db688e7d9ba30a4b4055e224c1e380348a3