Skill instructions
ci-pr-gate10.5 KB
View saved version →
---
name: ci-pr-gate
description: Turn a dependency change — a before/after package.json/lockfile snapshot, or one or more named "bump X from A to B" upgrades — into one deterministic PASS/WARN/FAIL verdict formatted for a CI check or PR-comment bot, not a conversational report. Use when the user explicitly wants a mergeable/blocking verdict, a "CI gate," a PR status comment, or asks "should this PR be blocked," "is this safe to merge," "gate this dependency bump." Not for a plain explanation of what changed in a diff (dependency-audit's own snapshot-diff flow already covers that conversationally) and not for a single already-installed package's trust investigation (package-trust-check).
---
# CI PR gate
Use this skill only when the user wants a **verdict**, not a report — the
output is meant to be pasted directly into a CI check or a PR comment by a
bot, so it has to resolve to exactly one of PASS / WARN / FAIL every time,
even on partial or ambiguous data. `dependency-audit`'s existing
"Comparing two snapshots" flow already produces a good conversational
explanation of a `diff_dependencies` result for a human reading it in
chat — don't duplicate that here. This skill's only job is applying a fixed,
documented policy on top of the same tool output to produce a scannable
verdict instead, and packaging it for a bot persona: no hedging, no
follow-up questions, no "as an AI" framing, no options to weigh — one
verdict, ranked blocking reasons, done.
If the user wants to investigate whether one already-installed package is
compromised (not a PR's dependency change), use `package-trust-check`. If
they want to know what to actually *do* about a finding, hand off to
`incident-response`.
## Determining input shape
1. **Two full snapshots** (any mix of `package.json`, `package-lock.json`,
`yarn.lock`, `pnpm-lock.yaml`) pasted as a before/after pair — call
`diff_dependencies({ before, after })` directly.
2. **One or more named bumps with no full snapshots** — the common
Renovate/Dependabot PR-title shape ("Bump lodash from 3.10.1 to
4.17.21," "upgrade minimist to 1.2.6") — call
`simulate_dependency_upgrade({ packageName, currentVersion,
targetVersion })` once per named package. There is no batch variant of
this tool. Cap it at 25 packages in one turn; if more were named, run
the first 25 and say explicitly which were skipped rather than silently
dropping them.
3. **Both** (full snapshots plus one or more specific packages the user
wants a deeper semver/breaking-change read on beyond what a diff
computes) — run `diff_dependencies` first, then add
`simulate_dependency_upgrade` calls only for the specifically-named
packages. Don't run both tools for the same package by default; that's
redundant.
4. **Nothing parseable** — ask for the before/after content or the exact
`name@current→target` bump(s). Don't guess a version that wasn't given.
## Steps
1. Route per the table above and make the tool call(s).
2. Collect every vulnerability finding worth ranking from the results:
any entry with `vulnerabilityDelta` of `"introduced"` or
`"still-vulnerable"` (from either tool), taking `id`/severity from its
`vulnerabilities[]` array (diff) or `targetVulnerabilities[]` (simulate).
Deduplicate identical `{packageName, cveId}` pairs, then call
`prioritize_remediation({ findings })` once with all of them — even for
a single finding, since KEV status isn't visible from the raw severity
string alone.
3. Apply the [Gate policy](#gate-policy) below to every package checked,
deterministically. The overall verdict is the worst single-package
verdict (FAIL beats WARN beats PASS) — one failing package fails the
whole gate.
4. Produce the [Output contract](#output-contract) below. That is the
entire response — no preamble, no "let me check that for you," no
summary paragraph after it.
## Gate policy
Evaluate every rule for every package checked; the highest tier any rule
triggers is that package's verdict.
### FAIL — block the merge
- `installScriptIntroduced === true` on any changed/upgraded package
(from either tool). This is the same highest-signal field
`dependency-audit` and `diff_dependencies`'s own description call out —
a routine-looking bump quietly adding a `postinstall` is the shape of a
compromised-maintainer attack, and a gate should never let that through
as a warning.
- Any `vulnerabilityDelta: "introduced"` finding whose `highestSeverity` /
vulnerability severity is `CRITICAL` or `HIGH` — **regardless of what
tier `prioritize_remediation` assigns it.** This was confirmed directly:
simulating minimist's 1.2.6→1.2.5 downgrade (which reintroduces the real
CRITICAL CVE-2021-44906) and ranking that finding through
`prioritize_remediation` returns `tier: "monitor"`, `score: 11.37` —
because EPSS's 30-day exploitation probability for that CVE is currently
only 4.6% and it isn't KEV-listed. `prioritize_remediation`'s tier
answers "what should I work through first across my whole backlog,"
which is the right question for a fix-priority ranking but the wrong one
for merge admission control — a PR that actively introduces a CRITICAL
vulnerability shouldn't pass just because that CVE isn't trending right
now. So: severity on an *introduced* finding is a hard block on its own;
`prioritize_remediation`'s tier is what decides WARN-level ordering
below it, not whether this rule fires at all.
- Any finding — introduced or pre-existing — that `prioritize_remediation`
ranks `tier: "patch-now"` (CISA KEV-listed, confirmed active
exploitation). This fires even at MEDIUM/LOW severity, same as that tool
documents: active exploitation overrides severity.
### WARN — pass, but flag for human review
- `vulnerabilityDelta: "still-vulnerable"` (pre-existing, not introduced by
this change) at any severity — real, but not this PR's fault; don't
block the PR for it, but don't hide it either.
- An introduced finding at MEDIUM/LOW severity, or any finding ranked
`tier: "patch-soon"`, `"scheduled"`, or `"monitor"` that didn't already
trigger a FAIL rule above.
- `riskTier: "breaking-change-likely"` or `isBreakingBySemver: true`
(simulate only) — not a security issue, but a gate consumer needs to
know a major/breaking bump is riding in.
- `targetDeprecated` set (simulate only — `diff_dependencies` doesn't
surface a per-package deprecation flag; that's a known gap in this
skill's coverage for the snapshot-diff path, not something to work
around with an extra tool call here).
- `engineChange.tightened === true` (simulate only) — the target version
now requires a newer Node than the current one supports.
- `targetIsPrerelease === true` (simulate only).
- Any unresolved side: `resolutionNote` set (diff) or
`currentVersionNote`/`targetVersionNote` set (simulate) — e.g. a
git/workspace/file specifier, or a target range with no satisfying
published version. Never treat unresolved as PASS; it means the gate
couldn't actually check anything for that entry.
- `changeType === "downgrade"` with no vulnerability reintroduced — still
worth a reviewer's eyes; a PR that quietly lowers a dependency version
for no stated reason is unusual enough to flag.
### PASS
Every package checked triggered none of the above.
## Output contract
Lead with one machine-parseable line, then a short human-readable body.
Nothing goes above the verdict line.
```
GATE: <PASS|WARN|FAIL>
## Dependency gate — <✅ PASS|⚠️ PASS WITH WARNINGS|❌ FAIL>
Checked N package(s), M flagged.
### Blocking (only when FAIL)
- `pkg@version`: <one-line reason — name the exact CVE/GHSA id if there is
one> ([npmscanUrl])
### Warnings (omit section if none)
- `pkg@version`: <one-line reason> ([npmscanUrl])
### All packages checked
| Package | Before → After | Verdict | Reason |
|---|---|---|---|
```
- The `GATE:` line's value must match the verdict implied by the rest of
the response — never let the table show a FAIL-tier row while the top
line says PASS.
- Every row's "Reason" cites the specific field that triggered it (exact
CVE/GHSA id and severity, or "installScriptIntroduced", or "major semver
bump," etc.) — never a bare "flagged," and never fold multiple reasons
for the same package into a vague summary when the row has more than
one; list them.
- Keep the table to the packages actually checked — don't restate
`removed` entries from a `diff_dependencies` result; removing a
dependency isn't a gate concern.
- A package the tool couldn't resolve still gets its own row with verdict
WARN and the resolution note as the reason — never drop it from the
table silently.
## Do not
- Do not derive the verdict from raw severity strings alone once
`prioritize_remediation` has run — its `tier` (not the bare severity) is
what decides WARN-vs-FAIL ordering among findings that didn't already
hit the severity-based FAIL rule above.
- Do not treat a `prioritize_remediation` tier below `patch-now` as
license to pass an *introduced* CRITICAL/HIGH finding — see the FAIL
rule above; that tool's tier is a fix-priority ranking across a whole
backlog, not a merge-admission signal on its own.
- Do not block a PR for a `still-vulnerable` (pre-existing) finding the PR
didn't introduce — WARN it, don't FAIL it.
- Do not run both `diff_dependencies` and `simulate_dependency_upgrade` on
the same package in the same request "just in case" — route per the
input-shape table once.
- Do not silently cap a >25-package batch of named bumps without saying
which ones were skipped.
- Do not add a conversational summary, caveats paragraph, or follow-up
question after the output contract — the structured verdict is the
entire deliverable. If something is genuinely too ambiguous to gate
(falls under "nothing parseable" in the input-shape table), say so
instead of producing a contract with a guessed verdict — don't emit a
fabricated PASS/WARN/FAIL over data you don't have.
- Do not imply this skill actually posts the comment, sets a commit
status, or blocks the merge itself — it only produces the text; whatever
bot or workflow is consuming this response is responsible for the actual
enforcement action.
- Do not silently drop an unresolved entry into the PASS bucket — every
unresolved side is a WARN with the resolution note as its reason, per
the Gate policy above.
## Tools used
`diff_dependencies`, `simulate_dependency_upgrade`, `prioritize_remediation`
— see `agents/openai.yaml` for the MCP server dependency, and
`references/test-prompts.md` for the test cases to run in ChatGPT
Developer Mode before submitting.
Referenced files: 2
dependency-audit15.4 KB
View saved version →
---
name: dependency-audit
description: Audit a project's npm dependencies for known vulnerabilities and risky install scripts before installing, upgrading, or shipping. Use when the user pastes or attaches a package.json/lockfile, lists dependencies, or asks to check/audit/scan their packages for security issues — including comparing two snapshots (a PR diff) or checking license compliance.
---
# Dependency audit
Use this skill when the user wants a security check across multiple npm
packages at once (a `package.json`, a lockfile, or a plain list of
`name@version` pairs) — not for a question about a single package. A plain
factual single-package question ("what does X do") the model can answer
directly with `get_package` or `query_vulnerabilities`; a single-package
*trust* question ("is X safe," "was X compromised/hijacked") should use the
`package-trust-check` skill instead, which runs the deeper
maintainer-history and publish-provenance checks this skill deliberately
reserves for already-flagged packages only. A forward-looking question about
what to *add* rather than what's already installed ("should we add X," "X
vs Y vs Z for this job," "what should we use to do X") should use the
`new-dependency-evaluation` skill instead — it orchestrates
`compare_packages`/`suggest_alternative` for exactly that decision, which
this skill's per-inventory vulnerability/license flow isn't built for.
## Input
Accept dependency name+version pairs from:
- Pasted `package.json` contents (use `dependencies`; only include
`devDependencies` if the user asks to include dev dependencies — say
explicitly which set you audited).
- Pasted lockfile contents.
- Pasted CycloneDX JSON or SPDX JSON SBOM contents.
- A plain list the user typed, e.g. "lodash 4.17.15, express 4.17.1".
- Raw `npm audit --json` output (npm 7+'s `{vulnerabilities: {...}}` format,
or legacy npm 6's `{advisories: {...}}`) — do NOT re-parse this by hand
into a `packages` list for `batch_query_vulnerabilities`. Call
`enrich_npm_audit({ content })` directly instead: it parses the report
itself, resolves each finding's GHSA advisory to a CVE alias via OSV
(npm audit JSON almost never carries a CVE id on its own), and returns the
same patch-now/patch-soon/scheduled/monitor ranking as step 9 below in one
call — this replaces steps 1, 3, and 9 for the findings it covers, though
steps 4-8 (install scripts, maintainer/provenance checks, license
compliance) still need the package names it flagged run through those
tools separately, since npm audit's own JSON has no license/maintainer/
install-script data. `yarn audit --json`/`pnpm audit --json` are not
supported — treat those like a lockfile paste (step 1) instead.
- A GitHub repository URL (e.g. "audit https://github.com/owner/repo") — do
NOT ask the user to paste `package.json`/lockfile contents in this case.
Call `audit_github_repository({ url })` directly instead: it fetches the
manifest/lockfile from the repo's default branch itself and runs steps 3,
4, 6, and 8 below (vulnerability, install-script signal, license
compliance, and — for whatever it flagged as critical/high severity, a
possible typosquat, or deprecated — the `check_maintainer_changes`/
`check_package_provenance` ownership checks too, up to 5 packages per call,
prioritized the same way step 6 already ranks them) in one call, replacing
that part of the flow below. Check `ownershipCheckNote` for any flagged
package past that 5-package cap and call the two ownership tools on those
directly. Still follow up per-package with `analyze_install_script` (for a
package the tool's lighter `installScriptScanScope:
"lifecycle-scripts-only"` signal flagged but didn't deep-scan — check
`deepScanNote`) and `suggest_alternative` the same way you would from the
pasted-content flow.
If the user instead pastes **two** snapshots and asks what changed (a PR,
before/after, "did this upgrade introduce anything") — see
[Comparing two snapshots](#comparing-two-snapshots-pr-review) below instead
of the single-inventory flow.
If the message contains no parseable package list at all, ask the user to
paste their `package.json` or lockfile — or give a GitHub repo URL — rather
than guessing at what to audit.
## Steps
1. Prefer passing the raw pasted inventory straight to
`batch_query_vulnerabilities` via its `content` input when the user gave a
`package.json`, lockfile, CycloneDX JSON, or SPDX JSON document. Only
manually build a `packages` array when the user gave a plain dependency
list instead.
2. For raw `package.json` content, set `includeDevDependencies: true` only if
the user explicitly asked to include dev dependencies; otherwise the tool
defaults to production-ish dependencies only. Be explicit in the answer
about what set you audited. For `yarn.lock`, note that the file itself
cannot distinguish production from dev dependencies, so the tool will scan
every resolved package and emit a warning about that.
3. Call `batch_query_vulnerabilities` once with the parsed/raw input. The tool
now chunks large inventories internally — do not re-chunk the request in
the skill layer unless the model/runtime itself forces a request-size limit.
Each finding already includes severity, a summary, CVE aliases, and the
fixed version — do not call `query_vulnerabilities` again per flagged
package just to re-fetch detail you already have. The only exception:
if the result has an `enrichmentNote` (a very large audit crossed the
enrichment cap), the vulnerabilities it names are ID-only — call
`query_vulnerabilities` on those *specific* packages if the user needs
full detail on them.
4. For every package the batch call flags, follow up with `get_package`
(or `get_package_version` when an exact version was provided) to check
maintainers, license, and install scripts (`preinstall`/`postinstall`).
Treat install scripts as a separate risk signal from known CVEs, not
something to fold into the same score. If the user wants to know what a
flagged install script actually *does* rather than just that one exists,
follow up with `analyze_install_script` — it fetches the published
tarball and statically scans the script and the files it references
against npmscan's red-flags rubric, returning a `totalScore`/`riskTier`.
Also surface what `get_package`
already computes for you: `deprecated`/`maintenanceSummary` (a
deprecated or abandoned dependency is a real finding, not just a CVE
footnote) and `possibleTyposquatOf` (if set, this package's name is one
typo away from a much more popular one — flag it prominently as a
supply-chain risk to verify, not as confirmed malice).
5. If the user wants coverage beyond the direct dependencies you were given
(asks about "transitive"/"indirect" risk, or the inventory is small — up
to 15 root packages), call `analyze_transitive_dependencies` instead of, or
in addition to, step 3. It walks each root's own dependency tree (default
depth 2, capped at 3) and returns `vulnerablePaths` naming which direct
dependency actually pulled in each vulnerable transitive package —
`batch_query_vulnerabilities` alone only ever checks the exact packages
listed. It has no `content` shortcut, so build the `packages` array by
hand from the parsed inventory.
6. For a package that's flagged as critical/high severity, deprecated, a
possible typosquat, or that the user specifically calls suspicious, add
the two ownership/supply-chain checks — don't run these for every clean
package in a large audit, they're expensive and only useful signal on
elevated-risk packages:
- `check_maintainer_changes` — reconstructs maintainer-add/remove history
from the npm packument and flags account-takeover patterns (a new
maintainer who published shortly after being added, a sudden full
maintainer-list replacement, a long-standing maintainer quietly
dropped) plus GitHub repo transfers/archival. If it flags a newly added
or fully turned-over maintainer, follow up with
`check_maintainer_blast_radius({ maintainerUsername })` on that
account — it lists every other package the same account currently
touches and flags a tight cluster of packages published within a short
window of each other, the compromised-account shape behind the 2025
chalk/debug ("qix") incident (~18 packages within ~2 hours). A large
total package count alone is not a red flag; only a tight cluster is.
- `check_package_provenance` — checks npm's Sigstore publish provenance
against reality: does the attested source repo/commit match
`package.json`'s declared repository, is this package missing
provenance while its npm-scope/maintainer peers consistently have it,
and does the tarball's install scripts/dependencies match what's
actually committed at the attested source commit (a mismatch here is
the stolen-npm-token publish pattern). Structural only, not a
cryptographic re-verification.
7. Only call `get_latest_advisories` if the user separately asks for
broader npm-ecosystem context — it is not part of the default flow.
8. If the user asks about license policy/compliance (or pastes an
allow/deny list), call `check_license_compliance` with the same parsed
package list. With no `policy` given it applies the default enterprise
rule (copyleft/network-copyleft/proprietary = violation); pass
`policy: { allow, deny }` when the user states their own rule. Report
`needsReview` licenses (unrecognized/mixed SPDX expressions) separately
from confirmed violations — don't silently treat "unknown" as compliant.
9. Once you have the full set of flagged CVE/GHSA findings (from step 3
and/or 5), and there is more than a couple of them, call
`prioritize_remediation` with one `{packageName, cveId, severity,
currentVersion, fixedVersion}` entry per finding to get a patch-now /
patch-soon / scheduled / monitor tier per finding (CISA KEV status
overrides everything else; EPSS exploitation probability is the primary
ranking signal otherwise; severity is the fallback). Lead the summary
report with this ranking instead of a flat severity list — it answers
"what do I fix first," which is usually what the user actually needs from
an audit with more than a few findings.
10. For any package that ends up flagged as deprecated, vulnerable at its
latest version, abandoned/stale, or a confirmed typosquat, offer (don't
force) a replacement: `suggest_alternative` combines the maintainer's own
deprecation hints with category-matched search results and returns
plain-language `whySuggested` notes per candidate, or a
`nonPackageAlternatives` entry when a language built-in supersedes the
package entirely. Call it when the user asks what to use instead, or
proactively name it as available in the summary rather than always
running it unasked for every flagged package.
11. Produce one summary report: a table of package → flagged issue(s) (CVE/
GHSA id + severity + fixed version, "risky install script", "deprecated/
unmaintained", "possible typosquat", "maintainer/provenance anomaly",
and/or "license violation") → the `npmscanUrl` from that tool's result →
a one-line recommendation (upgrade to the fixed version, patch, replace,
or no action needed). If `prioritize_remediation` ran, order the table by
its tier/rank instead of by package name.
## Comparing two snapshots (PR review)
When the user gives a **before** and **after** snapshot (any mix of
`package.json`, `package-lock.json`, `yarn.lock`, `pnpm-lock.yaml`) and asks
what changed, call `diff_dependencies({ before, after })` directly — this
replaces the batch-query flow above, it doesn't precede it.
- Lead with `installScriptIntroduced` findings — a routine-looking version
bump that quietly adds a postinstall script is the shape of a
compromised-maintainer supply-chain attack, and is the single highest-
signal field this tool returns.
- Report `vulnerabilityDelta` per changed package (introduced / fixed /
still-vulnerable / still-clean), not just a final isVulnerable flag — the
direction of the change is the point of a diff.
- Note the detected `beforeFormat`/`afterFormat` and any `comparisonNote`
when the two snapshots are different formats (e.g. a range in
`package.json` resolved against a pinned lockfile version).
- yarn.lock has no direct/transitive distinction, so a diff against a
yarn.lock covers every resolved package in the file, not just direct
dependencies — say so if relevant to what changed.
## Output requirements
- Always include the `npmscanUrl` for every flagged package so the user can
read the full write-up on npmscan.com.
- Keep known-vulnerability, install-script, maintenance/deprecation,
typosquat, maintainer/provenance, and license risk visually separate — a
package with no CVEs but a `postinstall` script (or a `possibleTyposquatOf`
flag) is not "clean."
- When a CVE has a `fixedVersion`, name it directly in the recommendation
("upgrade to X.Y.Z") rather than a generic "upgrade the package."
- State which set was audited (e.g. "checked 24 production dependencies,
skipped devDependencies").
- If the input was an SBOM, state that non-npm entries (if any) were skipped
and surface the tool's warning/ignored-count metadata when relevant.
## Do not
- Do not guess a version that wasn't provided — call `batch_query_vulnerabilities`
without a version rather than inventing one.
- Do not fabricate CVE/GHSA ids, severities, or fixed versions beyond what
the tools returned.
- Do not silently drop non-npm SBOM entries — say they were skipped because
npmscan's vulnerability pipeline is npm-only.
- Do not treat `possibleTyposquatOf` as proof of malice — it's a rule-based
heuristic (name similarity + low popularity), not a verdict. Report it as
"worth verifying," matching the tool's own hedged language.
- Do not run `check_maintainer_changes`/`check_package_provenance` across an
entire large inventory by default — reserve them for flagged/suspicious
packages or an explicit request, they're per-package deep checks, not a
batch scan.
- Do not treat a missing `provenance` or a repository transfer as confirmed
compromise on its own — both tools return hedged findings meant to prompt
verification, not a verdict.
- Do not treat a large `totalPackagesFound` from `check_maintainer_blast_radius`
as a red flag by itself — many legitimate maintainers publish hundreds of
packages over a career. Only a `tight-publish-cluster` finding (several
packages' latest versions landing within a short window of each other) is
the actual signal.
- Do not invent a replacement package name — `suggest_alternative` already
distinguishes maintainer-named replacements from category-matched guesses
and reports `nonPackageAlternatives` when no package is the right answer;
don't override that with your own guess.
## Tools used
`batch_query_vulnerabilities`, `get_package`, `get_package_version`,
`analyze_install_script`, `analyze_transitive_dependencies`,
`check_maintainer_changes`, `check_maintainer_blast_radius`,
`check_package_provenance`,
`check_license_compliance`, `diff_dependencies`, `prioritize_remediation`,
`enrich_npm_audit`,
`suggest_alternative`, `get_latest_advisories`, `audit_github_repository` —
see `agents/openai.yaml` for
the MCP server dependency, and `references/test-prompts.md` for the test
cases to run in ChatGPT Developer Mode before submitting.
Referenced files: 2
incident-response8.9 KB
View saved version →
---
name: incident-response
description: Turn a flagged npm supply-chain finding — or just a vague, non-technical description of one ("this package looks sketchy," "someone said we got hacked," "npm install did something weird") — into concrete remediation steps. Use whenever the user wants to know what to actually do about a suspicious/compromised/flagged package, however precisely or vaguely they describe it — not for the initial trust/vulnerability investigation itself (see package-trust-check/dependency-audit for that).
---
# Incident response
Use this skill once the user wants to know what to actually *do* about a
supply-chain concern, not just what was found — whether that concern is a
precise tool finding already in this conversation, or just a worried,
non-technical description of a symptom. Most users will not know terms like
"lifecycle script" or a `rule` id, and will not think to run an npmscan tool
themselves first — reformulating a vague description into the right call is
this skill's job, not something to push back on the user for. This is the
last step, not a replacement for `package-trust-check` (single-package trust
questions) or `dependency-audit` (multi-package audits): those skills
investigate and report; this one turns a concern — precise or vague — into
an actionable plan.
## Reading a vague or non-technical request
Do not require the user to already know a `rule` id, a playbook name, or
npmscan's own vocabulary. Translate what they actually said using their
intent, not their exact words. Common shapes and what they map to:
| What the user says (any rough equivalent) | What to do |
|---|---|
| Names a specific package and says it's "sketchy," "hacked," "compromised," or just "is this safe" | Run `package-trust-check`'s investigation first (get_package, check_maintainer_changes, check_package_provenance, analyze_install_script as applicable) — you need real findings before a remediation plan means anything. Once findings exist, continue to step 2 below. |
| Names a package AND describes a specific symptom ("it downloads something during install," "the maintainer changed") | Run the one finding tool that actually checks that symptom (analyze_install_script for install-time behavior, check_maintainer_changes for ownership, check_package_provenance for publish/build mismatches) rather than the full trust-check sweep — faster, and still grounds the answer in a real finding. |
| Describes a symptom with NO package name — "what do we do about a postinstall that downloads a binary," "how do we respond to a typosquat," "what's the process for a maintainer takeover" | No finding tool has anything to check. Go straight to `get_remediation_playbook({ id: "..." })` with your best-guess playbook id from the table below. A wrong guess is harmless — it comes back `matched: false` — so guess confidently rather than asking the user to be more specific first. |
| Totally generic, no package, no symptom — "we got a security alert, what do I do," "help, is this bad" | The one case worth a single clarifying question: ask what specifically happened or which package/alert, since there is nothing yet to ground a plan in. Don't guess a playbook id from nothing. |
Symptom → playbook `id` guesses for the no-package-name case:
- downloads/runs a binary, executable, or installer during install → `postinstall-binary`
- name looks like / is close to a well-known package, "typosquat," "fake version of X" → `suspected-typosquat`
- general "hacked," "compromised," "malicious code," "supply chain attack" with no more specific detail → `supply-chain-compromise`
- runs shell commands, `exec`, spawns a process during install → `child-process-in-install`
- "calls out," "connects to a server," "phones home" during install → `unexpected-network-install`
- "new maintainer," "ownership changed," "ownership transfer," "ex-employee's account" → `maintainer-change-flagged`
- "provenance," "build doesn't match the repo," "install script wasn't in the source," "someone bypassed CI" → `provenance-mismatch`
## Steps
1. Read the request per the table above. If a finding already exists in
this conversation (a prior `analyze_install_script`,
`check_maintainer_changes`, or `check_package_provenance` result), skip
straight to step 2 with its `findings[].rule` value(s).
2. Call `get_remediation_playbook({ rules: [...] })` with every distinct
`rule` value from the finding(s) in one call — it batches and deduplicates,
so don't call it once per rule. Use `id` instead for a symptom-only
request with no live finding (see the table above).
3. If the underlying finding's `riskTier` came back `none` (or the findings
array was empty), say plainly that no remediation is needed rather than
still calling this tool to manufacture a plan — an incident-response
playbook implies there's something to respond to.
4. For each `matched: true` result, present that playbook's `title`,
`severity`, and `steps` directly in your answer — the step `text` plus
its `why`, in order — not a paraphrase and not just a link. Lead with the
match's own `situationNote` when one exists (what this specific rule
actually found) before the steps, so the response is grounded in the real
finding rather than reading as generic boilerplate — this matters most
when a batch has several different rules landing on the same playbook
(see step 6). If you got here from an `id` guess rather than a real
finding (no `situationNote`), say so plainly ("this sounds like it might
be X — here's that playbook; tell me more if it's actually something
else") rather than presenting a guess as a confirmed diagnosis. Close
with the playbook's `preventionTips` when the user's question has any
forward-looking angle ("how do we stop this happening again," a policy/CI
question, or just naturally as part of a full incident report), and cite
its `references` (each labeled `incident` or `reading`) so the user can
verify the real-world grounding themselves — don't call an `incident`
reference a "reading" or vice versa, that distinction is deliberate.
Include the `npmscanUrl` too so the user can also read it on the docs
site.
5. For any `matched: false` result, say so plainly using the tool's own
`note` (e.g. "this finding is baseline/informational, not something with
its own response plan") — don't silently drop it or invent generic advice
in its place. If your own `id` guess came back unmatched, try the
next-closest guess from the table, or ask the one clarifying question
from the "totally generic" row rather than giving up.
6. When a findings array maps to more than one distinct playbook (e.g. a
`check_maintainer_changes` result with both a turnover finding and a
repository-transfer finding pointing at the same playbook, or a mixed
maintainer + provenance investigation pointing at two different ones),
present each matched playbook once, not once per originating finding.
## Do not
- Do not write your own remediation steps once a finding tool has run —
call `get_remediation_playbook` and use its actual `steps`/`why` text,
don't paraphrase or improvise around it.
- Do not refuse to help, or ask the user to look up a `rule`/playbook id
themselves, just because their request was vague — that is exactly the
case the table above exists for. Only ask a clarifying question in the
genuinely-nothing-to-go-on case.
- Do not treat `matched: false` as "this is safe" — it means no dedicated
playbook exists for that specific rule, which is a different statement
than "no risk here." Say what the tool's `note` actually says.
- Do not skip a finding tool and go straight to an `id` guess when the user
named a specific package — a guess is only for the no-package-name,
symptom-only case. When a package is named, ground the answer in a real
finding first.
- Do not present an `id`-guessed playbook as if it were a confirmed
diagnosis — say plainly that it's your best read of their description.
- Do not call `get_remediation_playbook` once per rule when a findings array
has several — batch all the rule values into one call's `rules` array.
- Do not force a playbook onto a clean/`none`-risk-tier result just because
the user asked "what should I do" — say no action is needed.
- Do not embellish a playbook's `references` beyond what's returned — cite
only the `label`/`url` the tool gave you, and preserve whether the tool
marked it `incident` (a real, verified compromise) or `reading`
(background material, not itself a breach) rather than upgrading a
`reading` reference into a claimed incident.
## Tools used
`get_remediation_playbook`, plus whichever of `analyze_install_script`,
`check_maintainer_changes`, `check_package_provenance` is needed to produce
the finding this skill responds to — see `agents/openai.yaml` for the MCP
server dependency, and `references/test-prompts.md` for the test cases to
run in ChatGPT Developer Mode before submitting.
Referenced files: 2
new-dependency-evaluation13 KB
View saved version →
---
name: new-dependency-evaluation
description: Help decide what to add as a NEW npm dependency — comparing 2-5 named candidates for the same job, evaluating one named candidate against its real peers, or shortlisting candidates from a described need, before anything is installed. Use for "which should we use for X," "axios vs got vs node-fetch," "is X a good pick for Y," "what should we use to do Z," "should we add X or is there something better." Not for auditing packages already in the project (see dependency-audit), not for a deep single-package compromise/trust investigation (see package-trust-check), and not for a plain factual question with no choice being made ("what does X do").
---
# New dependency evaluation
Use this skill when the user is making a **forward-looking choice** about
what to add to a project — not investigating something already installed.
It orchestrates two tools that already do the hard work individually:
`suggest_alternative` (turns one candidate name into a real shortlist of
peers) and `compare_packages` (fans out enrichment across 2-5 named
candidates and returns a structured side-by-side with a deterministic pick).
This skill's only job is routing to the right combination of the two based
on how many candidates the user actually named, then presenting the result.
If the user instead asks whether a package **already in their project** is
safe to keep, use `package-trust-check` (single package) or `dependency-audit`
(a pasted inventory) — those investigate what's already there; this skill
evaluates what to add next. If the question has no choice or decision angle
at all ("what does X do," "what's the latest version of X"), just answer
directly with `get_package` — don't invoke this skill's tool chain for that.
## Determining what to compare
Count how many named candidates the user actually gave, then route:
1. **2-5 named candidates for the same job** (e.g. "axios vs got vs
node-fetch," "should we use dayjs or date-fns") — go straight to
`compare_packages({ packages })`. This is the common case and needs no
extra tool call first.
2. **Exactly 1 named candidate** ("should we add uuid," "is left-pad a good
choice," "can we use X for Y") — a single package has nothing to be
compared against yet. Call `get_package({ name })` for the named
candidate's own `description` (`suggest_alternative`'s `source` object
doesn't include one), then `suggest_alternative({ name, reason:
"general", limit: 4 })` to generate real peers (never invent
plausible-sounding package names yourself). Screen the suggestions
against that description per
[Screening candidates before comparing](#screening-candidates-before-comparing-routing-cases-2-3)
below, then call `compare_packages` with `[name, ...screened
suggestions, up to 4]` so the original candidate is scored against real
competition instead of judged in isolation. If `suggestions` comes back
with fewer than 1 usable peer after screening (e.g. a built-in
replacement fully covers the case, like `left-pad` →
`String.prototype.padStart()`), skip `compare_packages` — there's
nothing left to compare — and report the `nonPackageAlternatives` plus
the `get_package` baseline on the named candidate instead.
3. **0 named candidates, just a described need** ("what should we use to
parse dates," "we need something for HTTP retries") — call
`search_packages({ query })` using the user's own description as the
query. Build a 2-5 name shortlist from the results, preferring the
highest `popularityTier`/`maintenanceTier` matches, skipping any result
carrying `possibleTyposquatOf`, and screening each result's
`description` against the user's *stated need* (their own words are the
comparison anchor here — no extra fetch needed) the same way as case 2
— then run `compare_packages` on that shortlist. If nothing in the
search results looks like a real contender (all very-low
popularity/stale, or nothing actually matches the described need), say
so rather than forcing a comparison, and ask the user for a starting
name or two.
4. **More than 5 named candidates** — `compare_packages` caps at 5. Don't
silently drop candidates without saying so: run it on the first 5 and
name which were left out, or ask the user to narrow the list if the
ones dropped seem like they'd matter to the decision.
### Screening candidates before comparing (routing cases 2-3)
`suggest_alternative`'s `categoryOverlap` and `search_packages`'s result
ordering are both keyword/token-overlap signals, not semantic relevance —
and a single shared generic word (e.g. both packages' descriptions mention
"guid") combined with high popularity can rank an unrelated package above
genuinely relevant ones. This was confirmed directly: comparing candidates
for `uuid` (an RFC9562 UUID *generator*, `description: "RFC9562 UUIDs"`)
surfaced `win-guid` (`description: "Windows legacy GUID parser"` — an
unrelated Windows binary-format tool) as the top suggestion, ranked above
the real UUID library `@paralleldrive/cuid2`, purely because `win-guid`
happens to have very high download counts and shares the single word
"guid." Reading the two descriptions side by side makes the mismatch
obvious immediately ("generates RFC-standard UUIDs" vs. "parses legacy
Windows binary GUID structures") in a way `categoryOverlap`'s token count
alone does not catch. So for every candidate before it enters
`compare_packages`:
- Actually read and compare full description text, not just token
overlap — the source's own `description` (from `get_package` in case 2,
or the user's stated need in case 3) against the candidate's
`description`. Ask in plain terms: does this candidate do the same job,
or does it just share vocabulary with something that does? Drop
candidates that fail this even if their `categoryOverlap` list looks
populated — shared words are a hint to go check, not a verdict on their
own.
- Treat a `categoryOverlap` of exactly one generic word (`"guid"`,
`"data"`, `"util"`, etc.) as weak supporting evidence at best; two or
more overlapping tokens, or overlap on a specific/technical term, is
stronger — but the description comparison above is the actual decision,
not the token count.
- Don't let a high `weeklyDownloads`/`popularityTier` on its own excuse a
poor description match — a package can be extremely popular as a
transitive dependency of something unrelated to the job at hand.
## Steps
1. Route per the table above and make the tool call(s).
2. Read `differentiators` before `recommendation` — it names which
candidate(s) stand out on downloads, GitHub stars, TypeScript support,
known vulnerabilities, deprecation, typosquat flag, and install-script
risk. This is what makes the comparison legible; don't just report the
final pick with no supporting detail.
3. Report `recommendation.pick`, `runnerUp`, and `rationale` verbatim —
don't substitute your own judgment for the deterministic score unless a
`candidates[]` entry shows something the score can't see (e.g. the user
already said they need TypeScript-first and two candidates are close).
Always state `confidence` too — a `"low"` confidence pick between two
close candidates is a materially different answer than a `"high"`
confidence one.
4. Surface every `found: false` candidate with its `resolutionError`
(typo? unpublished? malformed name?) rather than silently dropping it
from the comparison you present.
5. Note that `installScriptRisk` here is the **lifecycle-scripts-only**
signal (`scanScope: "lifecycle-scripts-only"`) — it scans the command
strings, not the tarball. If the recommended pick has a nonzero
`installScriptRisk.totalScore` and the user is about to actually install
it, mention that `analyze_install_script` (via `package-trust-check`)
gives the deeper, tarball-aware scan before they commit — don't run it
automatically as part of this skill, just point at it.
6. If the user's decision also turns on license policy (they mention a
license constraint, or ask "which of these is safe to use license-wise"),
follow up with `check_license_compliance` on the shortlist — it's not
part of the default flow, only pull it in when license is actually in
play.
7. Sanity-check `githubStars` against `weeklyDownloads` per candidate:
`githubStars` is attributed to whatever repository the candidate's own
`package.json` declares, unverified — a tiny, low-download package
showing a huge star count (e.g. thousands of stars on a package with
under a thousand weekly downloads) can mean its declared `repository`
field points at a different, unrelated project's repo rather than its
own (confirmed directly: `@aigne/uuid`, ~850 weekly downloads, declares
`repository: github.com/uuidjs/uuid` — the real `uuid` package's repo,
not its own — and so inherits that repo's star count). Flag a large
downloads/stars mismatch like this as worth independent verification
before adopting the package, rather than reporting the star count at
face value as a credibility signal; this skill's tools don't run the
deeper source-attestation check that would confirm or rule this out
(`check_package_provenance`, via `package-trust-check`).
8. Treat `downloadTrend.changePercent` with caution when
`weeklyDownloads` is low (roughly under a few thousand) — a small
absolute change produces a large, noisy percentage (e.g. a candidate
with 850 weekly downloads showing `"growing"` at 800%+ off a tiny prior
base is statistical noise, not real momentum). Lead with the absolute
download figure, not the percentage, for any low-volume candidate.
## Output requirements
- Always include each candidate's `npmscanUrl` so the user can read the full
write-up.
- Present the comparison as a table (or clearly separated per-candidate
summary) with at minimum: downloads/trend, popularity/maintenance tier,
deprecated status, latest-version vulnerability status, TypeScript
support, GitHub stars, and install-script risk tier — then the pick and
rationale below it, not interleaved.
- When `suggest_alternative` was used to generate the peer set (routing
case 2), say so explicitly ("compared against N real peers in the same
category, not just this one package in isolation") so the user
understands the comparison isn't limited to what they originally named.
- State plainly that this is a point-in-time snapshot (downloads,
vulnerabilities, and maintenance activity all change) — not a permanent
verdict, especially if the user's decision is time-sensitive.
## Do not
- Do not invent candidate package names yourself when the user gave 0 or 1
— `search_packages`/`suggest_alternative` exist specifically so the
shortlist is real, current registry data instead of names recalled from
training knowledge that may be renamed, abandoned, or gone since.
- Do not call `compare_packages` with fewer than 2 or more than 5 packages
— dedupe first (case-insensitive; the tool itself rejects exact
duplicates with a 400), and route through cases 2-4 above instead of
forcing a single name through it.
- Do not treat a deprecated or typosquat-flagged candidate as excluded from
the comparison — `compare_packages` still returns them (with `deprecated`/
`possibleTyposquatOf` set) so the user can see exactly why they lost; only
`recommendation.pick` is guaranteed to skip them, not the `candidates`
list itself.
- Do not silently narrow a >5-candidate list without telling the user which
names were dropped and why.
- Do not use this skill to re-investigate a package the user already has
installed and is worried about — that's `package-trust-check` or
`dependency-audit`; this skill's tools are tuned for choosing among
healthy-looking options, not for compromise/takeover forensics.
- Do not present `recommendation.pick` as a security clearance — it's a
weighted popularity/maintenance/vulnerability/typosquat/install-script
score, not a guarantee the package is free of issues the lighter scan
can't see.
- Do not pass every `suggest_alternative`/`search_packages` result straight
into `compare_packages` on the strength of `categoryOverlap` alone — a
single generic shared token plus high popularity can rank an unrelated
package first (verified: `win-guid`, a Windows GUID *parser*, outranked
a real UUID library when comparing alternatives to `uuid`). Read each
candidate's `description` and drop ones that aren't actually the same
tool for the job before comparing them.
- Do not report a candidate's `githubStars` as a plain credibility signal
without checking it against `weeklyDownloads` first — a low-download
package with implausibly high stars likely has a `repository` field
pointing at a different project's repo, not evidence of its own
popularity.
## Tools used
`compare_packages`, `suggest_alternative`, `search_packages`, `get_package`
(for the source description in routing case 2), and optionally
`check_license_compliance` — see `agents/openai.yaml` for the MCP server
dependency, and `references/test-prompts.md` for the test cases to run in
ChatGPT Developer Mode before submitting.
Referenced files: 2
package-trust-check6.92 KB
View saved version →
---
name: package-trust-check
description: Investigate whether one specific npm package is trustworthy — maintainer/ownership takeover signals, publish-provenance mismatches, and risky install scripts. Use when the user asks if a single named package is safe, compromised, hijacked, suspicious, or "can I trust this" — not for auditing a full package.json/lockfile (see dependency-audit) and not for a plain factual question like "what does X do."
---
# Package trust check
Use this skill when the user names **one specific package** (optionally a
version) and asks a trust question about it — "is X safe to use," "was X
compromised," "should I be worried about X," "check X's maintainers." This
is a deep, single-package investigation: it runs checks dependency-audit
deliberately skips for every package in a batch because they're too
expensive to run at scale.
If the user instead pastes a `package.json`/lockfile/dependency list or asks
to audit multiple packages, use the `dependency-audit` skill instead — don't
run this skill's checks across a whole inventory. If the question is purely
factual with no trust/safety angle ("what does lodash do," "what's the
latest version of express"), just answer directly with `get_package` — don't
invoke the full investigation for that. If the user is instead choosing
between candidates for something not yet installed ("should we add X," "X
vs Y for this job"), that's a forward-looking pick, not a trust
investigation — use the `new-dependency-evaluation` skill.
## Steps
1. Baseline with `get_package` (or `get_package_version` if the user gave an
exact version). Pull `deprecated`, `maintenanceSummary`,
`possibleTyposquatOf`, `isLatestVersionVulnerable`/`highestSeverity`,
`downloadTrend`, and days-since-last-publish. This alone answers a good
chunk of "should I trust this" and grounds the deeper checks that follow —
don't skip straight to the maintainer/provenance tools without it.
2. Call `check_maintainer_changes({ name })`. It reconstructs maintainer
add/remove history from the npm packument and flags:
- a maintainer added recently who then published shortly after (the
account-takeover pattern behind the Sept 2025 chalk/debug "qix"
compromise and ua-parser-js),
- a full sudden replacement of the maintainer list,
- a long-standing maintainer quietly dropped,
- a maintainer-list change on npm not yet tied to any release — flag this
as the *more* urgent case, since access changed hands but nothing has
shipped with it yet, so there's no version to warn the user off of.
Also reports GitHub repository transfer/archival — a transfer isn't
automatically hostile (e.g. jade → pug was a documented rename), say so
rather than treating every transfer as a red flag.
3. Call `check_package_provenance({ name, version })`. Three checks in one
call: does the Sigstore build attestation's source repo/commit match
`package.json`'s declared repository; is this version missing provenance
while its npm-scope or maintainer peers consistently publish with it (a
real anomaly, not just "no provenance" — plenty of legitimate packages
predate the feature entirely, which this tool already accounts for via
the peer baseline); and does the tarball's actual install scripts/
dependencies match what's committed at the attested source commit — a
script or dependency on npm that was never committed in source is the
stolen-npm-token publish pattern. This is structural verification, not a
cryptographic re-check of the Sigstore bundle — say so if the user asks
how deep it goes.
4. If `get_package`/`get_package_version` in step 1 showed
`hasLifecycleScripts`/a `preinstall`/`postinstall`/`prepare` entry, follow
up with `analyze_install_script({ name, version })`. It fetches the
published tarball and statically scans the script — and the files it
references — against npmscan's red-flags rubric (child_process,
network calls, `.ssh`/`.aws`/`.npmrc`/`*TOKEN`/`*KEY` access,
obfuscation, remote binaries off untrusted hosts, exfil endpoints,
eval-on-decoded-content), returning a `totalScore`/`riskTier`. A nonzero
score isn't automatically malicious — a legitimate binary download (e.g.
`cypress`) scores nonzero too — so report the actual `findings`, not just
the score.
5. If the combined picture ends up genuinely concerning (a maintainer-
takeover pattern, a provenance mismatch, a `possibleTyposquatOf` hit, or a
critical install-script finding), offer — don't force —
`suggest_alternative({ name, reason })` with `reason` set to whichever of
`"vulnerable"`/`"abandoned"`/`"typosquat"`/`"general"` best fits, so the
user has a next step instead of just a warning.
6. Produce one findings report, most-concerning signal first:
- State the verdict per check plainly (e.g. "no maintainer-change red
flags in the lookback window," not silence-as-clean).
- Every finding from `check_maintainer_changes`/`check_package_provenance`
is a hedged signal meant to prompt verification, not proof — say
"worth verifying independently," matching the tools' own language,
especially for anything below their `high`/`critical` tiers.
- Include the `npmscanUrl` so the user can read the full write-up.
- If nothing turned up anything but a low/none `riskTier` and no maintainer
or provenance findings, say so plainly — this skill exists to give
confident "looks clean" answers as often as it flags real risk.
## Do not
- Do not run this skill's checks across every package in a list — that's
`dependency-audit`'s job, and `check_maintainer_changes`/
`check_package_provenance` are deliberately reserved there for
already-flagged packages only, not a full inventory.
- Do not treat a missing `provenance` field, a repository transfer, or a
single maintainer addition as confirmed compromise on its own — report
what the tool actually flagged (or didn't) and let the `riskTier`/points
speak, don't editorialize past it.
- Do not skip `get_package` and jump straight to the deep checks — the
baseline (deprecated, popularity/maintenance tiers, typosquat) is cheap
and often already answers the question.
- Do not invent a replacement package name yourself — let
`suggest_alternative` return its own `whySuggested`/`nonPackageAlternatives`
rather than guessing.
- Do not silently skip `analyze_install_script` when lifecycle scripts exist
just because `check_maintainer_changes`/`check_package_provenance` came
back clean — a legitimate maintainer can still ship a genuinely risky
script, and vice versa; report all applicable signals, not just whichever
ran first.
## Tools used
`get_package`, `get_package_version`, `check_maintainer_changes`,
`check_package_provenance`, `analyze_install_script`, `suggest_alternative`
— see `agents/openai.yaml` for the MCP server dependency, and
`references/test-prompts.md` for the test cases to run in ChatGPT Developer
Mode before submitting.
Referenced files: 2