← Plugin catalog
Security

Tahr Security

Tahr Security Inc v0.3.3

Review applications with an evidence-first security workflow that maps attack surfaces, tests authentication and access controls, traces dangerous inputs, models threats, and verifies fixes.

Language: English · Automatically detected from descriptions.

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
GPL-3.0-only
Package author
Yack Security Inc

Declared capabilities

  • Read
  • Write

Package observed Sep 30, 2026.

Files & skills

File archives

Plugin package81 files · 347 KBBrowse files →
Skill instructions
tahr-audit-android5.36 KB

View saved version →

---
name: tahr-audit-android
description: Audit Android application security from an APK, AAB-derived APK, Android source repository, manifest, or authorized emulator/device. Use for mobile release reviews, OWASP MASVS-oriented assessments, exported component and deep-link testing, WebView and IPC review, local storage and token analysis, mobile API traffic review, runtime instrumentation, privacy testing, and Android hardening validation.
---

# Tahr Audit Android

Use an artifact-first workflow to turn manifest and code signals into focused runtime checks. Keep static evidence, runtime reachability, and verified security impact distinct.

## Set scope and safety

1. Identify the application ID, build variant, APK hash, source revision, device profile, identities, and backend environment in scope.
2. Treat a local repository or supplied artifact as authorized for read-only review. Default to static analysis when active device or backend testing is not clearly authorized.
3. Use a disposable emulator snapshot, test install, test accounts, and synthetic data for state-changing checks.
4. Do not clear package data, change protected account credentials, trigger lockouts, submit real payments, send messages, write through content providers, load persistent code, or tamper with production data.
5. Do not modify the app or implement fixes unless the user explicitly asks.

Use secrets, tokens, PII, keys, cookies, and credentials only transiently for authorized verification. Persist their class, source, key or field name, redacted excerpt, length, fingerprint, scope, expiry, and replay result—not raw values.

## Inventory before testing

Read [attack-surface.md](references/attack-surface.md) before exploring the application.

Prefer existing artifacts over broad rescans: manifest, decompiled source, smali/resources, class or string index, static findings, device information, traffic capture, dynamic plan, and earlier coverage. If no compact index exists, create a focused inventory of URLs, secrets, crypto APIs, log calls, password/token fields, WebViews, JavaScript bridges, dynamic loading, and native libraries.

Build a target list with provenance for:

- exported activities, services, receivers, providers, and their permissions;
- deep links, app links, intent actions, URI authorities, and parameters;
- WebViews, loaded origins, settings, file/content access, and bridges;
- authentication, biometric, token, OAuth, and session paths;
- storage locations, backup behavior, logs, clipboard, and caches;
- crypto operations, keys, IVs, RNG, signing, and integrity decisions;
- endpoints, trust configuration, cleartext paths, pinning, and PII flows;
- deserialization, reflection, commands, dynamic code, native libraries, and dependencies;
- privacy permissions, tracking SDKs, consent states, retention, and screenshots.

## Separate evidence levels

Classify each observation immediately:

1. **Static candidate** — source, smali, manifest, resource, dependency, or configuration evidence identifies a plausible weakness.
2. **Runtime reachability** — adb, UI, proxy, logcat, filesystem, Frida, or backend evidence shows that the path executes or is externally callable.
3. **Verified impact** — the behavior crosses a confidentiality, integrity, authorization, authentication, privacy, or protected-functionality boundary.

Never label a static flag as runtime exploitation. An exported component, weak algorithm, permissive WebView setting, cleartext allowance, dependency CVE, or disabled resilience control is a target until the matching proof gate is met.

## Plan and run focused checks

Read [proof-gates.md](references/proof-gates.md) before dynamic testing or severity assignment.

Prioritize high-impact targets from artifacts. For each target, define the safe command or interaction, caller and app state, expected secure behavior, impact proof, cleanup, and stop condition. Exercise fresh, logged-in, proxied, unproxied, and offline states only when they add relevant evidence.

Use read-only or marker-based component and provider checks by default. Instrument only the classes and methods identified by static evidence. Capture metadata rather than raw values. Stop when the device, app, backend, or account shows instability.

For dependency or platform intelligence, search only exact detected packages and versions. Record affected range, fixed version, prerequisites, and source, then prove that the vulnerable code path is packaged and reachable. Intelligence is a lead, never a finding.

When stuck, use a bounded fresh-agent pass to suggest missing test states or alternate proof paths. Treat suggestions as candidates and independently verify them.

## Report without losing uncertainty

Return three separate sections:

- **Verified findings:** include category proof, exact component or file, caller/app state, commands or interactions, before/after behavior, impact, and remediation.
- **Candidates:** include the static or runtime signal, missing proof, and safest next check.
- **Coverage:** mark every material target `tested`, `partially_tested`, `skipped_with_reason`, `inaccessible`, `blocked_by_environment`, or `unsafe_or_destructive_skip`.

Do not silently omit high-risk components because the emulator, login, proxy, root, Frida, or backend was unavailable. Do not interpret a blocker as evidence that the app is secure. Calibrate severity to demonstrated impact, not MASVS category names or scanner labels.

Referenced files: 3

tahr-audit-secrets-config7.05 KB

View saved version →

---
name: tahr-audit-secrets-config
description: Audit application-owned secrets, cryptography, dependency reachability, infrastructure-as-code, containers, CI/CD, cloud permissions, and runtime security configuration with evidence and false-positive controls. Use for repository hardening, deployment review, leaked-key triage, dependency/CVE review, exposed debug or admin surface checks, or pre-release configuration audits.
---

# Tahr Audit Secrets and Config

Find configuration and supply-chain weaknesses that become real attacker capabilities. Do not turn a keyword, permissive development setting, or package advisory into a vulnerability without proving production relevance and reachability.

## Establish scope and safety

Inspect source, committed configuration, lockfiles, build files, IaC, containers, CI/CD, and documentation read-only by default. Inspect git history only when the user includes it. Do not search unrelated home directories, credential stores, or `.git` working metadata.

Never validate a discovered credential against a live provider unless the user explicitly authorizes that exact action and the account is disposable. Never print, copy, commit, or persist a raw secret. Record type, file and line, source, scope, length, a short redacted preview, and a SHA-256 fingerprint.

Read [references/config-audit-matrix.md](references/config-audit-matrix.md) for the audit families. Read [references/exploitability-research.md](references/exploitability-research.md) before using advisory or CVE information.

## Inventory every configuration plane

Enumerate:

- application config and environment loading, including defaults and fallbacks;
- secret-manager, KMS, key-vault, certificate, and signing-key integrations;
- manifests and lockfiles for production, development, plugins, images, actions, and build tools;
- Dockerfiles, compose files, Kubernetes, Helm, Terraform, Pulumi, CloudFormation, Ansible, systemd, IIS, reverse proxies, and serverless definitions;
- CI workflows, release jobs, artifact publishing, package registries, caches, and deployment scripts;
- browser/mobile public configuration, source maps, runtime config endpoints, health, metrics, debug, admin, docs, and actuator surfaces;
- logging, telemetry, backups, data exports, error handling, and crash artifacts.

Mark generated, vendored, example, fixture, test-only, local-only, and production-relevant paths. Missing deployment material is a coverage gap, not proof of secure deployment.

## Review secret exposure

Search provider-specific formats and contextual assignments for cloud keys, OAuth clients, signing/session/webhook secrets, database and queue URLs, private keys, developer tokens, payment keys, AI provider keys, backup credentials, and privileged API keys.

For each hit:

1. Determine whether it is application-owned committed content, committed history, generated output, a placeholder, documentation, test fixture, environment reference, secret-manager lookup, public browser key, or local checkout metadata.
2. Exclude `.git/config`, remote URLs, hooks, logs, and scanner checkout credentials from application findings. A secret genuinely committed in repository history remains in scope.
3. Determine the exposure path: shipped client bundle, public artifact, image layer, CI log, repository audience, runtime endpoint, backup, or developer-only file.
4. Determine likely privilege, environment, restrictions, rotation status, and blast radius without using the raw value.
5. Inspect whether the application fails closed when the secret is absent or falls back to a default, empty, or hardcoded value.
6. Classify the result as `verified`, `static-confirmed`, `candidate`, or `rejected` with exact reasons.

Do not report environment-variable references or vault lookups as hardcoded secrets. Treat public client identifiers and publishable keys according to provider design; require missing restrictions or privileged use before claiming impact.

## Review production controls

Trace configuration from source default through environment override to deployed consumer. Look for:

- debug/test modes, stack traces, verbose errors, install/setup routes, sample accounts, default credentials, and unrestricted metrics or admin endpoints;
- fail-open authentication, authorization, origin, webhook-signature, feature-flag, or network-policy behavior when configuration is missing or malformed;
- wildcard or reflected credentialed CORS, unsafe cookie/session flags, proxy trust, generated-link host trust, weak TLS, missing transport enforcement, and cache-key confusion;
- overly broad IAM, public storage, unauthenticated services, exposed databases, unrestricted egress, cloud metadata access, and security-group/network-policy gaps;
- privileged containers, root users, dangerous capabilities, host namespaces, writable mounts, Docker socket exposure, unpinned images, and secrets embedded in layers;
- untrusted pull-request code reaching privileged CI secrets, mutable third-party actions, artifact poisoning, unsafe interpolation, and excessive workflow permissions;
- sensitive logging, personal data in URLs, backups without access controls, client-side secret storage, and long retention;
- weak password hashing, encryption, randomness, key/nonce/IV reuse, insecure verification, or custom cryptography tied to a real security property.

Read the code or deployment path that consumes each setting. A permissive example file or a development-only branch is not a production finding unless it can affect a shipped environment.

## Establish dependency reachability

For every dependency lead:

1. Prove the exact resolved version from a lockfile, image digest, installed artifact, or reproducible build evidence.
2. Identify the affected function, class, feature, plugin, image component, action, or transitive path.
3. Prove first-party use and a reachable entrypoint, job, parser, upload, request, build, or deployment path.
4. State the attacker-controlled input or prerequisite.
5. Inspect wrappers, disabled features, sandboxing, network controls, version backports, and vendor fixes.
6. Separate `version-candidate`, `reachable-candidate`, `attempted`, and `proven` states.

Do not report a CVE because a package name appears. If runtime reachability cannot be established, produce a precise upgrade/hygiene note or validation target rather than an exploitable finding.

## Challenge and report

For every accepted claim, name the exact failed control and seek the strongest contradiction. Distinguish severity from confidence. Link configuration primitives into an attack chain only when every hop is independently proven.

Report:

- evidence-backed findings with redacted proof, production relevance, affected capability, remediation target, and regression or policy test;
- candidates needing runtime, cloud, build, or owner confirmation;
- rejected hits with placeholder/test/dev/framework control evidence;
- coverage by configuration plane, including missing deployment artifacts and unreviewed environments.

Do not say an application or deployment is secure when production configuration, runtime identity, cloud state, or dependency reachability was unavailable.

Referenced files: 3

tahr-map-attack-surface4.77 KB

View saved version →

---
name: tahr-map-attack-surface
description: Map the real security-relevant surface of a web application or API from source, specifications, JavaScript, browser behavior, and authorized traffic. Use for pre-pentest reconnaissance, security-review scoping, hidden route or parameter discovery, undocumented API inventory, role-aware surface comparison, or judging whether an existing review actually covered the application.
---

# Map the Attack Surface

Build an evidence-backed inventory before testing vulnerabilities. Treat every discovered item as coverage evidence, not as a finding.

## Set the boundary

1. Identify the repository roots, application origins, API origins, environments, and supplied specifications or traffic captures.
2. Record the allowed runtime scope. Do not contact a live target unless the user supplied it or clearly authorized testing it.
3. Default to source-only analysis when authorization, credentials, or a runnable environment are absent.
4. Keep runtime activity read-only and low volume. Do not submit destructive forms, create real orders, send invitations, modify accounts, or enumerate unrelated infrastructure.
5. Never print or persist passwords, cookies, bearer tokens, API keys, reset links, private keys, CSRF values, or user PII. Retain names, locations, value classes, lengths, and short SHA-256 fingerprints when useful.

## Inventory independent evidence sources

Inspect each available source independently before merging:

- server routes, controllers, RPC handlers, middleware, authorization declarations, background jobs, queues, and WebSocket/SSE handlers;
- OpenAPI, Swagger, GraphQL schemas and documents, generated clients, protobufs, and API examples;
- frontend routes, forms, fetch/axios clients, lazy chunks, source maps, feature flags, upload configuration, storage keys, and custom headers;
- infrastructure routes, reverse-proxy rules, serverless functions, storage buckets, callback handlers, and public documentation;
- authorized unauthenticated and authenticated browser/API traffic, separated by identity, role, and tenant.

Do not infer reachability from a route name alone. Mark each operation as source-derived, documented, browser-observed, traffic-observed, or runtime-confirmed.

## Build canonical operations

Read [surface-inventory.md](references/surface-inventory.md) and create one record per meaningful operation. Preserve:

- protocol, origin, method or operation type, path, content type, and request shape;
- query, path, body, header, cookie, form, multipart, and GraphQL variable fields;
- authentication state, identity/role/tenant context, object identifiers, owner hints, and workflow state;
- response class, state-changing behavior, source artifact, and confidence.

Do not merge operations when method, body shape, parser, auth state, role, tenant, owner, response behavior, or workflow state differs. Preserve exact endpoint-to-field provenance; a global parameter list is insufficient.

## Prioritize attacker-relevant surface

Rank concrete operations higher when they expose:

- login, reset, MFA, invite, token, session, or account lifecycle behavior;
- admin, role, tenant, membership, ownership, billing, entitlement, approval, or settings functions;
- object IDs in any carrier, bulk/composite IDs, exports, downloads, or sensitive response fields;
- upload, import, preview, render, conversion, callback, webhook, URL-fetch, or integration behavior;
- price, total, quantity, discount, status, state, owner, tenant, role, or idempotency fields;
- HTML/Markdown/template input, search/filter/sort expressions, file paths, XML, or serialized data;
- undocumented versions, hidden client routes, debug endpoints, source-map leads, or role-specific discrepancies.

Priority controls review order only. Do not discard lower-ranked operations.

## Close coverage gaps

Read [coverage-gates.md](references/coverage-gates.md). Assign every discovered item one terminal state: `visited`, `runtime_confirmed`, `source_only`, `queued`, `partially_explored`, or `skipped_with_reason`.

When discovery is unexpectedly shallow, perform one bounded second pass using a different evidence source: follow lazy routes, inspect request builders, parse schemas, expand safe UI elements, or compare another supplied identity. Never describe absent evidence as proof that a feature is absent.

## Deliver the inventory

Report:

1. scope and evidence sources inspected;
2. canonical operations grouped by trust boundary and identity context;
3. high-value targets with exact provenance;
4. identity, tenant, object, upload, callback, and workflow maps;
5. coverage gaps, blocked areas, and the consequence of each gap;
6. candidate hypotheses clearly labeled as unverified.

Do not assign vulnerability severity from recon alone. Recommend the appropriate Tahr testing skill for each candidate class.

Referenced files: 3

tahr-review-tahr-findings3.83 KB

View saved version →

---
name: tahr-review-tahr-findings
description: Read applications, assessments, and findings from an already configured Tahr MCP connection. Trigger only when the user explicitly asks to query, list, summarize, or review Tahr account data; do not trigger for generic security reviews, source-code reviews, or non-Tahr findings.
---

# Review Tahr Findings

Use this optional, read-only workflow for existing Tahr customers who have manually configured the Tahr MCP server in Codex. This is a Codex/local manual integration only. Never claim that it provides public ChatGPT account linking.

## Establish account context

Start every explicit Tahr-data request with `get_context`. Before querying account data, confirm and report the authenticated organization name and the relevant capabilities. If the user named a different organization, stop and ask them to switch or reconfigure the connection; do not query the authenticated organization.

Respect the reported capabilities:

- Access applications only when `canReadApplications` is true.
- Access findings only when `canReadFindings` is true.
- Treat this workflow as read-only even when `canEditFindings` is true.

Use only these tools: `get_context`, `list_applications`, `get_application`, `list_assessments`, `list_findings`, and `get_finding`. Never invoke mutation tools or perform write, triage, status, or comment actions.

## Handle connection failures safely

If the MCP server or tools are unavailable, authentication returns 401, the token is missing, revoked, or expired, permission fails, or the service is unavailable or rate limited, give concise setup or recovery guidance. Suggest checking that the personal token environment variable and Codex MCP configuration are present, restarting Codex, obtaining the required organization access, or retrying later as applicable. Never ask the user to paste a token into chat. Offer to continue with the independent local security skills.

## Resolve and query records

Resolve a human application name with `list_applications`, then use the exact application ID returned by Tahr. Do not reveal details for null, missing, or cross-organization resources.

List tools accept optional `cursor` and `limit` parameters. Use a limit from 1 through 50; the default is 25. Responses contain `items` and `page.{nextCursor,isDone}`. Paginate only when the user requests all or complete results. For each subsequent page, keep every filter unchanged and pass the prior `nextCursor`. Otherwise, stop when the request is satisfied and disclose the coverage limit.

For `list_findings`:

- Always provide `kind` as `security` or `authorization`.
- Use the default `assessmentScope: latest` unless the user explicitly requests history or all assessments; only then use `assessmentScope: all`.
- Filter as needed by `applicationId`, `assessmentId`, or severity: `Critical`, `High`, `Medium`, `Low`, or `Info`.
- When generic "findings" clearly means both security and authorization findings, query each kind separately. Otherwise, clarify the ambiguity before querying.

Use `get_finding` only when requested or needed for the requested detail. Treat returned records as Tahr platform records, not independently verified vulnerabilities. Preserve finding IDs and assessment IDs in summaries where helpful, and distinguish recorded evidence from claims.

If list items are empty, report only that no matching records were returned; do not speculate. If a get operation returns null, report that the resource is unavailable without disclosing whether it exists elsewhere.

## Protect data and report results

Do not send repository content, secrets, credentials, tokens, or unrelated chat data to Tahr. Never log or persist the bearer token.

Provide concise result summaries grouped by severity or application as requested. State whether the summary covers all matching pages or only the pages and item limit queried.

Referenced files: 1

tahr-secure-app7.76 KB

View saved version →

---
name: tahr-secure-app
description: Perform an evidence-backed, pentester-style security review of an application from source, configuration, specifications, tests, and optionally an explicitly authorized local or staging runtime. Use for comprehensive app security audits, pentest readiness, pre-release reviews, dangerous-flaw discovery, or coordinating the Tahr specialist skills; also use when a prior scanner or LLM review created confidence that needs independent verification.
---

# Tahr Secure App

Review the application as an attacker would: map what is reachable, form target-specific hypotheses, try to disprove each candidate, require exploit-class proof, and account for what was not tested.

## Set the review mode

Choose one mode and state it before reviewing:

- `deep`: inspect every admitted first-party file and close every high-risk gap. Use by default for “secure this app.”
- `focused`: review a named feature or boundary deeply and list excluded areas.
- `retest`: verify an accepted fix with `$tahr-verify-security-fix`.

Treat source and local artifact inspection as read-only review. Run runtime probes only against a local, disposable, or explicitly authorized test target. Do not infer permission to test production, third parties, other tenants, or real accounts. Use low-impact canaries, disposable objects, bounded concurrency, and reversible actions. Never persist raw passwords, tokens, cookies, private keys, personal data, or cloud credentials.

## Create the evidence ledger

Read [references/evidence-contract.md](references/evidence-contract.md) before classifying any issue. Use [assets/review-report.template.json](assets/review-report.template.json) when a durable JSON artifact is useful. Write artifacts only where the user requests or in a clearly named local review directory; do not modify application code during a review.

Keep three lanes separate throughout the work:

1. `findings`: claims that meet the applicable proof status.
2. `candidates`: concrete leads whose required proof is incomplete.
3. `coverage`: reviewed, partial, blocked, skipped, and not-applicable scope.

Never promote a scanner hit, dangerous function name, package advisory, route name, response status, reflection, timing change, accepted upload, or prior report statement by itself.

## Map before judging

Inventory the application before searching for bugs:

- first-party files, frameworks, manifests, lockfiles, infrastructure, deployment and CI configuration;
- HTTP routes, GraphQL resolvers, RPC/gRPC handlers, WebSockets/SSE, webhooks, uploads, callbacks, serverless functions, CLI entrypoints, queues, jobs, and externally influenced schedulers;
- actors, roles, service principals, tenants, organizations, groups, ownership fields, sessions, and recovery flows;
- assets, data stores, secrets, billing or entitlement state, admin/debug functions, AI models, RAG sources, tools, and external integrations;
- browser routes, lazy chunks, source maps, API clients, mobile deep links, and undocumented or legacy API versions.

Follow thin controllers into middleware, policies, services, repositories, serializers, templates, and asynchronous consumers. Mark each admitted item `reviewed`, `partial`, `blocked`, `skipped-with-reason`, or `not-applicable`. Low apparent risk is a reason to review briefly, not to disappear the file from coverage.

Use `$tahr-map-attack-surface` when the reachable surface or identity coverage is unclear.

## Build attacker hypotheses

Create target-specific hypotheses in four forms:

- `actor -> action -> resource -> owner/tenant boundary -> expected decision`;
- `untrusted source -> transformations -> dangerous sink -> expected control`;
- `workflow state -> attempted transition/replay/race -> invariant -> authoritative readback`;
- `deployment input/default -> privileged capability or sensitive asset -> compensating control`.

Prioritize unauthenticated paths, cross-user or cross-tenant boundaries, state-changing functions, sensitive exports, callbacks and outbound fetches, file processing, server-side rendering, privileged fields, legacy versions, recovery paths, background workers, and AI tool use.

## Route to specialist skills

Read [references/review-routing.md](references/review-routing.md) and invoke only the applicable specialist skills. A comprehensive review normally includes:

- `$tahr-test-authentication` for login, recovery, MFA, OAuth/OIDC, tokens, cookies, and session lifecycle;
- `$tahr-test-access-control` for a complete actor-resource-action model,
  path-specific authorization traces, proof-gated findings, and safe validation;
- `$tahr-trace-dangerous-inputs` for injection, browser sinks, outbound requests, parsers, files, and uploads;
- `$tahr-test-business-workflows` for state, pricing, quota, invitation, approval, replay, and race abuse;
- `$tahr-audit-secrets-config` for secrets, cryptography, dependencies, infrastructure, and deployment defaults;
- `$tahr-test-ai-agents` or `$tahr-audit-android` when those technologies exist;
- `$tahr-threat-model-app` when a full-system threat model and validation plan
  are requested for the entire existing application.

## Investigate and challenge

For every candidate:

1. Cite the exact route, file, line or symbol, actor, input, object, or configuration involved.
2. Trace reachability across files and processes; distinguish dead, test, generated, dependency, and production code.
3. Name the expected control and inspect its actual placement and applicability.
4. Search for the strongest contradiction: middleware, policy, tenant filter, validator, encoder, allowlist, safe parser, parameter binding, environment guard, or framework behavior.
5. Record the contradiction verdict as `not-contradicted`, `contradicted`, `partially-contradicted`, or `insufficient-evidence`.
6. Define the exact claim and proof needed before attempting runtime validation.
7. Establish a normal baseline and an expected-denial or benign negative control.
8. Use the smallest safe proof. Record request/state/browser/callback/readback evidence without raw secrets.
9. Search one bounded set of sibling routes, helpers, models, and variants after a strong seed; give each variant its own evidence.

Operational failure is a limitation, not target-side proof. Stale authentication, missing roles, WAF or rate limiting, callback outage, tool error, or target instability must not become either a finding or a false-positive conclusion.

## Close coverage and chain impact

Before concluding:

- resolve every high-risk candidate as accepted, rejected with exact control evidence, out of scope, or deferred with a specific blocker;
- list unreviewed files, endpoints, identities, tenants, workflows, sinks, environments, and runtime-only claims;
- distinguish “tested with no issue observed” from “not tested”;
- link only independently proven primitives into attack chains, and require evidence for every hop;
- avoid “secure” or “no vulnerabilities” language when material gaps remain.

A clean result means only: no verified findings were produced within the stated, completed scope. It is not a guarantee about omitted or blocked scope.

## Deliver the review

Lead with dangerous, reachable issues. For each finding include the exact claim, affected location, preconditions, required proof, positive and negative evidence, contradiction result, impact, confidence, limitations, remediation target, and a repo-native regression test. Keep severity separate from confidence.

Summarize candidates and coverage gaps after findings. Redact sensitive substrings while preserving type, source, length, scope, expiry where relevant, and a non-secret fingerprint.

Validate a JSON review artifact with:

```bash
python3 <skill-directory>/scripts/validate_review.py path/to/tahr-review.json
```

Use `--strict` only when claiming a deep review has no unresolved high-risk scope.

Referenced files: 5

tahr-test-access-control13.4 KB

View saved version →

---
name: tahr-test-access-control
description: Perform complete or focused, evidence-backed access-control review from source and optionally an explicitly authorized local or staging runtime. Model subjects, roles, tenants, resources, actions, properties, policy rules, enforcement points, and owner-attributed test cases; trace object-, function-, property-, role-, and tenant-level authorization through REST, GraphQL, web, job, and asynchronous paths; safely validate IDOR/BOLA/BFLA, mass assignment, privilege escalation, and cross-tenant isolation; and reject status-code or guessed-ID false positives. Use for authorization code review, multi-user or multi-tenant assessments, admin and role boundary analysis, pre-pentest review, or validation of a suspected access-control finding.
---

# Tahr Test Access Control

Determine exactly who can perform which action on which resource, property, or
function—and prove when the implementation violates that application-specific
rule. Produce an authorization model and proof ledger, not a status-code diff.

## Load the operating contract

Before reviewing:

1. Read [full-review-workflow.md](references/full-review-workflow.md) for scope,
   execution order, completion, and lifecycle rules.
2. Read [access-control-data-contract.md](references/access-control-data-contract.md)
   for stable IDs, exact enums, and record relationships.
3. Read [access-matrix.md](references/access-matrix.md) before constructing
   operation, relationship, carrier, property, or variant coverage.
4. Read [access-proof-gates.md](references/access-proof-gates.md) before
   accepting or rejecting a candidate.
5. Read [runtime-test-safety.md](references/runtime-test-safety.md) before any
   runtime action.
6. Read [source-review-patterns.md](references/source-review-patterns.md) when
   source is available.
7. Use [worked-example.md](references/worked-example.md) only when the expected
   evidence-to-finding trace is unclear.

Start from [access-control-review.template.json](assets/access-control-review.template.json)
and keep [access-control-review.schema.json](assets/access-control-review.schema.json)
as the canonical output contract. Maintain one `access-control-review.json`;
derive all reader-facing artifacts from it.

## Choose scope and assurance honestly

Set `review_mode` to:

- `full` for every admitted authorization-relevant operation in the existing
  application; or
- `focused` for explicitly named operations, findings, resources, or policy
  boundaries.

Default to `full` when the user asks to review or secure the application and
does not explicitly narrow the authorization scope.

Set `analysis_basis` to `source_only`, `runtime_only`, or `hybrid`. A focused
review must carry a visible limitation and must not make an application-wide
claim. A source-only review may be `complete` with `source_observed` assurance
and planned runtime tests when every declared source surface is dispositioned.
It must not claim that an attack ran or that deployed enforcement failed.

Keep these states separate:

- `review_status`: whether the declared authorization scope was dispositioned;
- `assurance_status`: source observation versus authorized runtime validation;
- candidate `disposition`: lead, follow-up, rejected, source-confirmed, or
  runtime-confirmed;
- test `execution_status`: planned, passed, failed, inconclusive, or blocked;
- coverage `status`: reviewed, tested, pending, deferred, or out of scope.

Default to read-only analysis. Never start an application, send a request,
refresh a session, or mutate state unless the exact target and action class are
authorized. Never commit, patch, or reconfigure the reviewed application as
part of this skill.

## Freeze the review inputs

For source or hybrid review, create a deterministic manifest in the selected
output directory:

```bash
python3 <skill-directory>/scripts/build_review_manifest.py \
  path/to/application --include . \
  --output path/to/output/repository-manifest.json
```

Pass `--revision` for an immutable VCS or release revision. When omitted, the
script derives `snapshot-sha256:<digest>` from admitted paths and bytes. Copy
its embedded revision and content hash into metadata, manifest evidence, and
coverage inventory. Repeat `--package` for package/module labels and
`--document` for supplied repository policy or design files; documents are
admitted and hashed automatically. Use explicit empty arrays for absent
documents or exclusions. When the output lives under the application root,
place it in a dedicated subdirectory such as `.tahr-review/`; the builder
records and excludes that whole directory to prevent generated artifacts from
contaminating later snapshots, and refuses output directly in the root.

Inventory every admitted REST route, GraphQL query/mutation/subscription,
server action, RPC method, UI-backed function, webhook, worker/job, queue
consumer, export/download, bulk operation, legacy/versioned interface, and
administrative surface that makes or depends on an authorization decision.
Record intentionally public and non-applicable surfaces instead of deleting
them from the inventory.

When `$tahr-map-attack-surface` is available, consume its frozen operation
inventory and reconcile it; otherwise inventory locally. Companion skills are
optional—the access-control skill must remain independently usable.

## Establish identity and object truth

Model unauthenticated, user, peer, role, tenant, administrator, support,
service, worker, integration, and other applicable subjects. Keep source-modeled
identities separate from runtime-verified sessions.

For runtime identities, bind the redacted auth artifact fingerprint to an
observed caller, role, tenant or authorization domain, transport, freshness,
and validation evidence. A filename, configured label, JWT claim, profile
field, or successful HTTP response alone does not make a session trustworthy.
Mark unusable or mismatched identities as coverage limitations; never turn
their failures into target-side denials.

For each target object, record resource type, exact identifier carrier, owner
or controller, tenant/domain, sensitivity, lifecycle, and provenance. A shared
ID pool, guessed adjacent ID, globally discovered identifier, or caller-owned
`/me` object cannot prove peer or cross-tenant impact. Use owner-attributed
source evidence, an authoritative owner baseline, or a disposable object
created and read back under the owner identity.

Treat supplied assessment identities as protected. Do not delete, disable,
lock, re-role, rename, reset, or rotate them. Use fresh disposable objects and
non-protected accounts for state-changing validation.

## Build the authorization model before judging

Record subject, action, resource, property, relationship, tenant/domain,
workflow state, feature/plan, authentication strength, and other policy
context. Express each rule as `allow`, `deny`, `conditional`, or `unknown` and
cite its authority.

Prefer source policy, middleware, domain rules, role matrices, documented
requirements, or explicit assessment context. UI hiding, endpoint names,
generic assumptions about administrators, and an owner-success baseline may
corroborate a rule but do not normally prove expected denial alone.

For every operation, enumerate each independent authorization obligation:
source object, destination object, parent, child, relationship object,
property, function, and asynchronous continuation. Record every caller-supplied
identifier and sensitive property, where restrictions enter the path, where
they are consumed, the authoritative query or state change, and every
downstream enforcement point checked.

## Trace controls end to end

When source is available, trace shipped entrypoint to final data return,
mutation, worker, integration, signed URL, audit record, notification, or other
side effect. A route-level guard that permits an action somewhere is not proof
that the submitted target is in scope. A policy helper that exists but is not
consumed on this path is not a control.

Search for the strongest contradiction before retaining a gap: global
middleware, dependency injection, decorators, domain policy, repository
filters, ORM scopes, serializers, workers, database policy, deployment
controls, and sibling route variants. Mark dead, test-only, generated,
dependency, or unreachable paths explicitly.

Group operations only when policy, resource, action, enforcement point, and
relevant carriers and variants genuinely match. Preserve operation-level IDs
so grouping cannot hide a legacy route, bulk path, alternate parser, nested
GraphQL resolver, or asynchronous continuation.

## Construct complete matrix coverage

For each applicable operation, freeze the expected callers, relationships,
identifier/property carriers, and interface variants before recording results.
Include unauthenticated, own-object, same-role peer, cross-role, cross-tenant,
service-to-user, and privileged-function cases when the model makes them
meaningful.

Pair unauthorized cases with an authorized baseline of the same operation
shape. Change one declared authorization dimension at a time. If an identity,
owner-attributed object, parser, route variant, or safe fixture is unavailable,
keep the matrix cell and mark the exact blocker; never omit it to improve
coverage.

Ranking controls execution order, not inventory. A full review processes every
expected case or gives a specific disposition. Planned runtime validation does
not by itself make completed source coverage incomplete; missing high-risk
source analysis or an explicitly required runtime case does.

## Execute only safe, discriminating tests

Follow [runtime-test-safety.md](references/runtime-test-safety.md). Require the
test to match exactly one structured authorization target across origin,
environment, surface/operation/resource IDs, tenant or domain, identities,
transport, action and mutation scope, request/attempt limits, and validity
window. Use synthetic data and the smallest reversible proof.

Every test defines and distinguishes:

- caller-identity proof;
- owner/tenant/target attribution;
- authorized baseline success;
- expected denial or control-held signal;
- unauthorized protected-data/action impact or control-failure signal;
- authoritative readback and cleanup for state changes.

An executed result uses one `run_id`. Its freshly collected identity preflights
and every observed signal must cite same-run, same-target runtime evidence.
Keep purpose-specific evidence for each signal; one aggregate record cannot be
the sole proof for caller, target, baseline, denial/impact, readback, and
cleanup. `planned` and `blocked` tests never contain a result.

HTTP 200, non-empty output, size/hash differences, empty/null/false output,
generic SPA shells, validation errors, 4xx/5xx reachability, 202 acceptance, or
an echoed request are leads only. A mutation requires persistent authoritative
readback or an equivalent side effect. Repeat an accepted runtime failure with
fresh identity state and a fresh authorized control.

## Apply proof and false-positive gates

Use the five mandatory gates: caller, target, ownership or tenant, expected
denial, and unauthorized impact. A source-confirmed candidate must connect a
shipped reachable entrypoint to a concrete protected data/action sink, show the
missing or bypassed path-specific control, and record the strongest
contradiction checked. A runtime-confirmed candidate must additionally cite a
conclusive authorized test result with direct runtime evidence.

Reject or retain as follow-up:

- current-user endpoints returning only caller-owned or empty state;
- legitimate sharing, public, support, or administrator behavior supported by
  an applicable policy;
- guessed or shared IDs without owner attribution;
- wrong, stale, mismatched, or transport-incompatible identity material;
- soft 404s, shells, parser errors, resource absence, or unproven server errors;
- writes without readback, async acceptance without a side effect, or results
  whose baseline changed more than the authorization dimension.

Record rejected leads with the applicable control or contradiction evidence.
Do not erase them; they demonstrate that the review challenged its own leads.

## Challenge and publish the model

After the primary pass, use an independent subagent when available; otherwise
perform a separate adversarial reasoning pass with fresh instructions. Require
it to seek omitted operations and identities, unmodeled parent/child objects,
unused restrictions, alternate routes/parsers/versions, false owner or tenant
attribution, ambiguous collaboration, unsafe tests, unsupported impact, stale
evidence, and hidden coverage gaps. Record every challenge and disposition.

Validate the canonical model:

```bash
python3 <skill-directory>/scripts/validate_access_control_review.py \
  path/to/access-control-review.json --strict
```

Fix failures, rerun the challenger when material content changes, and render:

```bash
python3 <skill-directory>/scripts/render_access_control_review.py \
  path/to/access-control-review.json --output-dir path/to/output --strict
```

The renderer produces `access-control-review.md`, `validation-plan.json`,
`findings.json`, and `coverage.json`. Do not edit derived artifacts as separate
sources of truth.

Lead with confirmed source or runtime findings, unresolved high-risk
candidates, blocked tests, and coverage limitations. Never conclude that
access control or the application is secure. A clean result means only that no
additional proof-gated finding was produced within the declared completed
scope and assurance level.

Referenced files: 13

tahr-test-ai-agents4.98 KB

View saved version →

---
name: tahr-test-ai-agents
description: Test security boundaries in applications that use LLM chat, RAG or vector retrieval, memory, file or URL ingestion, model-rendered output, tool/function calling, MCP, or autonomous agents. Use for source-backed AI feature reviews, authorized local or staging runtime tests, prompt-injection assessments, cross-tenant retrieval checks, agent/tool abuse reviews, and AI resource-control testing.
---

# Tahr Test AI Agents

Review the application as an attacker crossing data, instruction, identity, rendering, and tool boundaries. Treat payload catalogs as aids; make the proof discipline the center of the assessment.

## Establish the boundary

1. Confirm the repository, feature, identities, and runtime target placed in scope.
2. Default to source-only analysis when authorization for active runtime testing is unclear.
3. Identify protected accounts, production data, third-party integrations, and actions that must remain read-only.
4. Do not modify application code, configuration, or deployed state unless the user separately requests remediation.
5. Use disposable tenants, documents, objects, tools, callback collectors, and marker values for active tests.

Never delete data, transfer value, send real messages, publish content, rotate credentials, change privileges, exhaust a customer budget, or exfiltrate real private data. Bound concurrency, output length, request counts, and cost.

## Build the real attack surface

Read [methodology.md](references/methodology.md) before selecting tests.

Trace model and embedding SDK calls to their actual HTTP, WebSocket, queue, or background-job entry points. Include supporting routes for uploads, knowledge bases, conversations, memory, feedback, model settings, tools, and shared views. When source is incomplete, inspect OpenAPI, client bundles, forms, runtime requests, streaming frames, and error shapes.

Confirm an endpoint only when it accepts AI-shaped input or produces generated, streaming, retrieval, embedding, model, or tool behavior. A route name containing `chat`, a generic JSON response, or a successful status code is not confirmation.

For each confirmed surface, record:

- endpoint, method, controllable field, and authenticated identity;
- tenant, conversation, memory, and retrieval context;
- model/provider and guardrail hints, marked `unknown` when unproven;
- ingestion formats and retrieval sources;
- renderer and downstream output sinks;
- available tools, resources, prompts, and autonomous steps;
- measured rate, token, output, streaming, and quota behavior;
- evidence source and confidence.

## Test conditionally

Establish a benign baseline before attack probes. Run only families whose preconditions exist:

- test direct instruction override against an application-specific control;
- test indirect injection only through content the application really ingests;
- test retrieval and memory across validated user or tenant boundaries;
- test output injection in the actual browser or downstream consumer;
- test tool and MCP agency at real resource and authorization boundaries;
- test resource controls with conservative measured probes;
- attempt chains only from previously observed signals.

Vary semantic attack concepts before cosmetic encodings. Record every attempt by endpoint, field, concept, payload hash, response class, and success criterion. For stochastic behavior, reproduce the same concept and boundary at least three times in five attempts unless the original security contract requires a stronger threshold.

When blocked, ask at most a few fresh-agent passes for concise new concept axes or plausible chains. Treat that advice as leads only. Never let an adviser, model claim, or payload classification verify a finding.

## Apply proof gates

Read [proof-gates.md](references/proof-gates.md) before promoting or scoring a result.

Maintain three separate collections:

1. **Verified findings** — the required boundary and impact proof exists.
2. **Candidates** — a concrete signal exists, but state the missing proof and safest next check.
3. **Coverage** — list tested, partially tested, skipped, inaccessible, degraded, and unsafe-to-test surfaces with reasons.

Do not promote model claims, marker-only obedience, generic policy text, stored-but-unretrieved content, API reflection, tool names, version intelligence, one lucky response, or theoretical cost.

## Handle evidence safely

Use raw secrets or private data only transiently when authorized and necessary. Persist the value class, source, tenant or role context, redacted excerpt, length, fingerprint, endpoint, payload class, and reproduction count. Do not place raw tokens, cookies, system prompts, private documents, credentials, PII, or tool output into findings, screenshots, logs, or reports.

Name findings after the demonstrated result, not the attempted technique. Calibrate severity to the proven data, action, resource, tenant, or browser boundary. End with prioritized remediation tied to the actual trust boundary and a concise residual-risk statement.

Referenced files: 3

tahr-test-authentication4.84 KB

View saved version →

---
name: tahr-test-authentication
description: Review and safely test web authentication and session boundaries across login, registration, password reset, magic links, MFA or OTP, OAuth/OIDC, SAML, passkeys, tokens, cookies, logout, and recovery. Use for authentication code review, pre-release auth testing, account-takeover analysis, session-management review, SSO integration review, or validating an existing security assessment.
---

# Test Authentication

Find identity-boundary failures, not merely unusual responses. Treat auth material as untrusted until it proves the intended identity and transport.

## Establish safe scope

1. Identify source roots, runtime origins, supported authentication methods, supplied identities, and expected account lifecycle.
2. Do not send runtime requests unless the user supplied or authorized the target. Use source-only analysis otherwise.
3. Treat all supplied accounts as protected unless explicitly labeled disposable. Do not lock, reset, disable, delete, re-role, enroll or remove MFA/passkeys, rotate credentials, or invalidate all sessions on protected accounts.
4. Use invalid identifiers for low-volume response-shape checks and disposable accounts for lockout, reset completion, password changes, MFA mutation, code replay, and takeover proof.
5. Stop runtime testing on lockout text, CAPTCHA, rate limiting, disabled-account state, unexpected notification delivery, or unclear side effects.

## Prove identity truth first

For each supplied identity:

- observe the rendered login flow before submitting credentials;
- classify password, split-step, OTP, MFA, magic-link, OAuth/OIDC/SAML, passkey, browser-bound, and custom stages;
- identify hidden state, nonce, CSRF, tenant, organization, provider, or login-method choices;
- validate success against an authenticated-only or identity-confirming endpoint;
- record the observed user, role, tenant, auth mode, and whether cookies/tokens are portable or browser-bound.

Do not equate a cookie, token, callback URL, HTTP 200, account picker, application shell, or pending MFA page with successful authentication. A failed role-specific login is a coverage blocker, not target access denial.

## Model the lifecycle

Trace these state transitions when present:

`registration/invite -> verification -> login -> step-up/MFA -> session refresh -> logout/revocation`

`forgot-password -> delivery -> token/code validation -> password change -> prior-session behavior`

`OAuth/SAML/passkey initiation -> provider/authenticator -> callback/completion -> application session`

Record every endpoint, browser action, actor, token class, binding, one-time expectation, expiry, and alternate/mobile/legacy channel. Derive endpoints from source, specifications, JavaScript, and observed traffic before using fallback names.

## Execute the abuse matrix

Read [auth-matrix.md](references/auth-matrix.md). Prioritize tests that cross an identity boundary:

- valid-disposable versus invalid account enumeration controls;
- rate limiting and weaker alternate endpoints;
- pre-auth, post-password, post-MFA, and fully authenticated stage skipping;
- token/code replay, wrong-account binding, stale-token reuse, and parallel requests;
- recovery, factor, password, email, and security-setting changes without reauthentication;
- session fixation, logout/timeout invalidation, CSRF, cookie scope, and refresh rotation;
- OAuth redirect/state/nonce/PKCE/code/client/scope binding;
- credentials or reusable secrets in URLs, responses, logs, JavaScript, caches, or browser storage.

Change one dimension at a time and pair every abuse attempt with a valid control. Refresh or re-establish the exact identity after an intentionally invalidating test before interpreting later responses.

## Apply proof gates

Read [auth-proof-gates.md](references/auth-proof-gates.md). Keep endpoint discovery, header observations, raw tokens, configuration smells, status codes, timing, and script labels in a candidate ledger until the class-specific proof gate passes.

For every confirmed issue, preserve:

- exact flow stage, endpoint/action, and actor/session context;
- baseline and manipulated request or browser action;
- token/cookie class and state using redacted fingerprints;
- authenticated-only data/action, wrong-account binding, replay, persistent state change, or other concrete impact;
- safe, faithful reproduction steps and cleanup or restoration status.

Never persist raw passwords, cookies, bearer/refresh tokens, authorization codes, reset/magic links, OTPs, SAML assertions, passkey material, secrets, or PII.

## Report coverage honestly

Separate `confirmed`, `candidate`, `not_reproduced`, `blocked_for_safety`, and `not_applicable`. List every untested lifecycle stage, missing disposable identity, browser-bound limitation, unavailable delivery channel, stale session, or provider blocker. Do not turn incomplete auth coverage into “no issue found.”

Referenced files: 3

tahr-test-business-workflows5.66 KB

View saved version →

---
name: tahr-test-business-workflows
description: Model and safely abuse-test stateful business workflows, API operations, and application invariants such as checkout, billing, credits, invitations, approvals, entitlements, exports, uploads, integrations, quotas, and asynchronous jobs. Use for business-logic review, race-condition and replay testing, mass-assignment or excessive-property review, workflow bypass analysis, API version/parser comparison, or pre-pentest testing of critical product flows.
---

# Test Business Workflows

Look for normal-looking requests that produce outcomes the business never intended. Prove the outcome, not merely that an unusual value was accepted.

## Bound the work safely

1. Identify source roots, specifications, runtime scope, supplied identities, sandbox/test mode, and critical workflows.
2. Do not send live requests without clear authorization. Use source, tests, schemas, and captured traffic to model workflows otherwise.
3. Treat supplied accounts, customer data, billing objects, orders, balances, coupons, invitations, and integrations as protected.
4. Require a disposable fixture or explicit dry-run/preview/test mode before creating orders, changing plans or roles, transferring value, consuming benefits, sending messages, registering webhooks, or mutating persistent state.
5. Define before state, expected effect, authoritative readback, restoration, and cleanup before every mutation or race batch. Stop on unexpected side effects or instability.

## Reconstruct the real workflow

For each critical workflow, combine code, UI, schemas, tests, background jobs, JavaScript, and authorized traffic. Record:

- actors, roles, tenants, objects, and ownership;
- entry conditions and server-side prerequisites;
- states, transitions, terminal states, and asynchronous steps;
- authoritative values and where each value originates;
- one-time tokens, idempotency keys, approvals, expiries, counters, and quotas;
- compensating actions, cancellation, rollback, and cleanup;
- alternate API versions, clients, content types, bulk operations, and direct endpoints.

Write explicit invariants such as “the server calculates the total,” “only the current owner may approve,” “a code is redeemed once,” or “a transition requires the immediately preceding state.” Cite the evidence for each invariant; do not invent product rules from route names.

## Derive an abuse plan

Read [workflow-abuse-matrix.md](references/workflow-abuse-matrix.md). Cover applicable classes:

- boundary values, type confusion, omitted fields, extra fields, nested/array forms, and duplicate parameters;
- client-controlled price, total, discount, tax, quantity, balance, owner, tenant, role, status, approval, entitlement, or feature fields;
- direct access to later steps, skipped/reordered prerequisites, repeated actions, stale token/state reuse, and alternate actors;
- duplicate submission, idempotency collisions, last-byte/single-packet concurrency in a disposable environment, and time-of-check/time-of-use gaps;
- rate/quota limits on sensitive successful operations;
- version, method, parser/content-type, mobile/legacy, bulk, GraphQL alias/batch, and asynchronous variants;
- excessive response properties, unsafe third-party responses, webhook signature/event handling, and generated-link/cache-key trust.

Rank high-impact invariants first, but give every discovered workflow or operation a terminal disposition. Do not replace target-derived requests with guessed endpoints or generic bodies.

## Execute paired experiments

For authorized runtime work:

1. Capture a valid single-request baseline and define what proves business success.
2. Change one invariant-related dimension while holding actor, object, method, body, parser, and timing constant.
3. Use a negative control and, where useful, an authorized control.
4. Read the authoritative state after the attempt; also inspect related list, audit, balance, entitlement, or downstream job state.
5. For races, first prove sequential behavior, then run the smallest bounded synchronized batch and count business successes—not merely HTTP responses.
6. Restore or delete only disposable fixtures and record cleanup.

An accepted value, 2xx, redirect, response-size change, missing header, or multiple concurrent responses is candidate evidence until the unsafe business outcome is shown.

## Apply proof gates and chain impact

Read [workflow-proof-gates.md](references/workflow-proof-gates.md). Confirm only reproducible outcomes such as unauthorized state transition, financial/value manipulation, duplicate benefit, quota bypass, stale-token replay, persisted privileged property, sensitive overexposure, weaker old-version control, or parser-dependent security bypass.

Cross-reference a proven workflow issue with access control and dangerous-input findings. Ask whether a single-object issue scales to bulk impact, a skipped approval unlocks a privileged action, or an upload/callback field reaches another trust boundary. Test one safe higher-impact hop only when authorized; do not inflate hypothetical chains.

## Report outcomes and gaps

For each finding, preserve the intended invariant, actor/object context, baseline, manipulated request/action, before/after proof, concrete impact, repeatability, cleanup, and faithful reproduction steps. Redact tokens, cookies, payment/customer records, PII, and secrets while retaining field names, value classes, lengths, hashes, and fingerprints.

Separate confirmed findings, candidates, expected behavior, unsafe variants not attempted, missing disposable fixtures, untested race conditions, unavailable roles, asynchronous jobs not observed, and other coverage gaps. Never describe incomplete workflow coverage as secure.

Referenced files: 3

tahr-threat-model-app12.4 KB

View saved version →

---
name: tahr-threat-model-app
description: Build a full, implementation-backed threat model of an entire existing application, covering actors, assets, trust boundaries, entrypoints, hop-level data flows, abuse cases, connected attack paths, security invariants, control gaps, risk responses, and executable validation handoffs. Use for comprehensive system threat modeling, security architecture assessment, pentest preparation, or correlating a complete application repository with configuration, IaC, API schemas, diagrams, and deployment documentation. Do not use for a feature-only, diff-only, or design-only review.
---

# Tahr Threat Model App

Model the entire existing application as an attacker would. Use source and
configuration for observed implementation, documents for intended behavior,
and runtime evidence only when the exact target and test are authorized.
Produce a decision and validation plan, not a code-smell list and not a
verified-vulnerability report.

## Load the operating contract

Before modeling:

1. Read [full-review-workflow.md](references/full-review-workflow.md) for the
   end-to-end sequence and completion gates.
2. Read [threat-model-ledgers.md](references/threat-model-ledgers.md) for the
   canonical record relationships and exact enums.
3. Read
   [threat-evidence-and-quality-gates.md](references/threat-evidence-and-quality-gates.md)
   before accepting threats, risk ratings, or a complete status.
4. Read [specialist-handoffs.md](references/specialist-handoffs.md) before
   assigning validation work to another Tahr skill.
5. Use [worked-example.md](references/worked-example.md) only when the expected
   evidence-to-test trace is unclear.

Use [threat-model.template.json](assets/threat-model.template.json) as the
starting structure and [threat-model.schema.json](assets/threat-model.schema.json)
as the output contract. Do not invent a different report structure.

## Establish full-application scope

Record the repository path and revision, packages, services, clients,
deployment environments, supplied specifications and documents, excluded
third-party internals, runtime authorization, previous model, and unanswered
questions. Keep `mode` equal to `full`.

Cover every admitted first-party application component. If time, access, or
missing evidence prevents full coverage, preserve the complete inventory,
disposition each gap, and set `model_status` to
`incomplete_high_risk_coverage`. Never silently narrow a full review.

Keep these concepts separate:

- `model_status`: whether the declared full scope has been dispositioned;
- `assurance_status`: whether conclusions are source-observed or also
  runtime-validated;
- `risk`: plausible impact and likelihood of a modeled threat;
- `confidence`: strength and completeness of supporting evidence;
- `execution_status`: whether a validation test is only planned or has an
  authorized result.

A source-observed model may be `complete` while validation tests remain
`planned`, provided runtime validation was not part of the declared scope and
all implementation evidence was dispositioned. Express the limitation through
`assurance_status`; do not misuse coverage status to imply a test ran.

Default to read-only analysis. Do not test production, third parties, real
tenants, or real accounts without explicit authorization. Redact credentials,
tokens, keys, personal data, and customer data from every artifact.

## Inventory before judging

Build a coverage baseline from all first-party files and supplied evidence.
Inspect source, manifests and lockfiles, configuration, CI/CD, IaC, containers,
API and GraphQL schemas, database models, tests, diagrams, role matrices,
workflows, integration notes, and deployment documents. Map:

- human, service, worker, administrator, support, peer, tenant, and third-party
  actors;
- critical business, identity, authorization, financial, operational, privacy,
  audit, and secret assets;
- clients, APIs, services, workers, queues, data stores, caches, renderers,
  control planes, AI systems, and external integrations;
- public, internal, administrative, legacy, debug, webhook, job, CLI,
  upload/download, import/export, socket, serverless, and mobile entrypoints;
- authentication, session, authorization, owner/tenant policy, validation,
  serialization, secrets, logging, rate, quota, and recovery controls;
- deployment, network, process, tenant, role, provider, browser, device, and
  asynchronous trust boundaries.

Group routes and files into capabilities and business workflows. Retain exact
locations as evidence, but do not turn the report into a route-by-route review.
Account for each high-signal item in `coverage`.

Freeze a deterministic repository manifest before review. Bind its embedded
content hash to the immutable revision and reconcile its admitted paths, packages,
environments, documents, and exclusions with metadata. Populate
`coverage.inventory.expected_subject_ids` from that inventory before changing
any coverage item from `pending`; never shrink the expected set to make a
review pass.

Create the manifest in the selected threat-model output directory:

```bash
python3 <skill-directory>/scripts/build_repository_manifest.py \
  path/to/application --include . \
  --output path/to/output/repository-manifest.json
```

Add repeated `--include` and `--exclude` arguments when scope is more precise.
Pass `--revision` for an immutable VCS/release revision; when omitted, the
script derives `snapshot-sha256:<digest>` from the admitted bytes. Copy its
embedded revision and content hash (also printed by the script) into metadata,
`coverage.inventory.manifest`, and manifest evidence. Use explicit empty arrays
for documents or exclusions when there are none; do not omit those scope fields.

When the reachable surface is non-trivial and `$tahr-map-attack-surface` is
available, send it the frozen revision and scope before deriving threats, then
reconcile its inventory into this model rather than treating its output as a
second source of truth. If that companion skill is unavailable, perform the
same stable inventory locally using this section; do not block or narrow the
review.

## Build an evidence-backed graph

Record material facts as claim-level evidence. Use exactly `observed`,
`intended`, `inferred`, or `unknown`; never combine classes in one field.

For every sensitive or state-changing flow, model each hop from the initiating
actor to the final response, durable state, or side effect. At each hop record:

- source, destination, protocol, input, and affected assets;
- actor, user, service, owner, tenant, role, and policy context;
- trust boundary crossed;
- validation, serialization, authentication, authorization, and logging
  controls;
- queues, workers, callbacks, redirects, repositories, providers, tools,
  browsers, and other continuation points.

Do not stop at a controller when another component performs the security
decision. Include data creation, replication, retention, deletion, backup,
residency, and third-party handling where material.

## Derive material threats

For each critical asset and flow:

1. State the security invariant.
2. Define a realistic actor, goal, preconditions, boundary, abuse steps,
   affected assets, and business impact.
3. Locate observed and intended controls and their enforcement points.
4. Search for the strongest contradiction in middleware, policy, service,
   repository, serializer, validator, framework, IaC, or deployment controls.
5. Reject, narrow, or mark the threat `validation_required` according to the
   contradiction result.
6. Rank risk with an explicit rationale and keep confidence separate.
7. Assign a response, decision, owner, next action, residual risk, and
   validation test.

Use STRIDE as a completeness prompt, not as evidence or a requirement to emit
one threat per category. Use ASVS, OWASP API, GraphQL, privacy, mobile, or AI
taxonomies only to find omissions and map controls. Retain only threats that
connect to the implementation-backed graph.

Build attack paths only from connected model IDs. Mark uncertain steps
conditional; do not invent a hop merely to make a chain more severe.

Keep the canonical model proportional to the application. Reuse a control,
invariant, decision, or validation test across related threats when the
enforcement point, owner, and discriminating oracle are genuinely the same.
Keep claim statements short and reference stable IDs instead of copying the
same narrative. Do not create reciprocal records solely to make the artifact
look complete.

If no candidate survives the evidence, contradiction, and materiality gates,
leave `threats`, `attack_paths`, `decisions`, `validation_tests`, and
`questions` empty. Preserve the populated inventory, graph, implemented
invariants and controls, evidence, coverage, and independent challenge. Never
invent a low-value threat or test to avoid an empty ledger.

## Add applicable specialist analysis

When personal or regulated data is present, trace collection, linkability,
identifiability, detectability, disclosure, consent/awareness, retention,
deletion, residency, and third-party processing.

When LLMs, agents, RAG, embeddings, MCP, prompts, memory, providers, or tools
are present, trace prompt injection, retrieval and memory isolation, tool
authorization, confused-deputy paths, output trust, provider disclosure,
supply-chain changes, evaluation bypass, and wallet/quota abuse.

When a lane is not applicable, record `applicable: false` with evidence. Do not
invent threats merely to populate a taxonomy.

## Create executable validation handoffs

Create a validation test for every material uncertain, missing, or potentially
bypassable control. Include the target Tahr skill, authorization required,
safe environment, fixtures, preconditions, normal baseline, exact action,
expected control, `attacker_case.attacker_success_signal` and
`expected_denial_signal`, `control_case.control_success_signal` and
`control_failure_signal`, evidence, cleanup, destructive risk, confidence, and
current execution status.

Treat each test as `planned` until authorized evidence proves otherwise. Never
convert a proposed test into a finding. Route the test to the specialist named
in [specialist-handoffs.md](references/specialist-handoffs.md).

## Challenge before publishing

Run a separate adversarial quality pass after producing the draft and before
publishing it. Use an independent subagent when available; otherwise use a
fresh, explicitly separate review pass. Give the reviewer the draft model and
source evidence, not the desired conclusions.

Require the reviewer to challenge missing assets, actors, boundaries, flow
hops, contradictory controls, unsupported impact, inflated risk, route-review
drift, unhandled documentation, incomplete coverage, weak actions, broken
references, and non-executable tests. Record findings and their dispositions in
`quality_review.challenge_findings`.
Resolve every high-severity review finding or keep the review and model failed.
The reviewer must not silently rewrite the model it is judging.

## Validate and render

For a durable model, write only to a user-selected output directory or a
clearly named `tahr-threat-model-output/` directory. Never modify application
source during the review.

Validate before presenting the model as final:

```bash
python3 <skill-directory>/scripts/validate_threat_model.py \
  path/to/threat-model.json --strict
```

Fix validation failures, rerun the independent challenge when material model
content changes, and validate again. Then render concise views:

```bash
python3 <skill-directory>/scripts/render_threat_model.py \
  path/to/threat-model.json --output-dir path/to/output --strict
```

The renderer produces `threat-model.md`, `validation-plan.json`, and
`coverage.json` from the canonical model. Do not edit derived files as though
they were independent sources of truth.

## Deliver concise-first results

Lead with:

1. model and assurance status;
2. the five most important connected attack paths or threats;
3. blocking security decisions and their owners;
4. the first five validation tests;
5. high-risk coverage gaps.

Place complete ledgers after that summary or in the canonical JSON artifact.
Avoid repeating the same threat narrative in every section.

Never conclude that the application is secure. A clean full model means only
that no additional material modeled risks were identified within the stated,
completed evidence scope. Preserve assumptions, limitations, accepted risks,
pending tests, and a review date so future changes can update the model rather
than starting over.

Referenced files: 11

tahr-trace-dangerous-inputs5.87 KB

View saved version →

---
name: tahr-trace-dangerous-inputs
description: Trace attacker-controlled input through parsing, validation, normalization, storage, and dangerous server or browser sinks, then safely validate exploitability with class-specific proof gates. Use for injection review, source-to-sink analysis, XSS, SQL/NoSQL injection, command or template injection, SSRF, XXE, path traversal, unsafe deserialization, file upload/processing, webhook, CORS/postMessage, or client-side trust-boundary testing.
---

# Trace Dangerous Inputs

Find complete attacker-controlled data paths. Do not report a dangerous API call, suspicious regex match, reflection, or error without proving reachability and impact.

## Set a safe review mode

1. Identify repository roots, runtime targets, specifications, traffic, authentication context, and authorized scope.
2. Default to static tracing when live testing is not explicitly authorized.
3. Keep runtime probes low volume, non-destructive, and tied to exact discovered operations. Do not test login credential fields, customer objects, real payment/order flows, or unrelated infrastructure.
4. Require a disposable fixture before uploads, persistent content, webhook registration, or other mutations. Define readback and cleanup first.
5. Use harmless unique markers, controlled callbacks, safe owned canaries, and low-impact commands only. Never delete data, establish persistence, dump broad files/databases, scan internal networks, or alter cloud resources.

## Build the source-to-sink map

Read [source-sink-matrix.md](references/source-sink-matrix.md). For every candidate path, record:

- entrypoint and exact source: path, query, body, nested field, header, cookie, form, multipart metadata, GraphQL variable, message, file, stored value, or browser source;
- parsing and canonicalization order, including decoding, type coercion, content-type selection, duplicate parameters, archive/document parsing, and redirects;
- validation, allowlisting, authorization, normalization, encoding, parameterization, and sanitization controls;
- transformations and trust-boundary hops across services, queues, jobs, databases, caches, templates, browsers, and third parties;
- final sink, execution context, and output/render/fetch/readback path.

Trace second-order behavior: input stored now may later reach a query, template, browser render, document processor, shell, webhook, or background job.

## Prioritize real sinks

Prioritize operations evidenced by code, schemas, JavaScript, or traffic:

- query builders and raw SQL/NoSQL/search expressions;
- shell/process APIs, dynamic evaluation, deserializers, and template engines;
- URL fetchers, redirects, webhooks, integrations, image/document renderers, and import/export jobs;
- XML parsers, file/path/archive operations, upload pipelines, and served-content behavior;
- HTML/DOM/navigation/eval-like browser sinks, postMessage handlers, and cross-origin data access.

Do not send a payload to the application root merely because a field name looks interesting. Preserve the actual method, body, parser, auth/role/tenant context, and provenance.

## Validate progressively

For an authorized runtime target:

1. Establish a clean baseline and the operation's real success semantics.
2. Send a unique inert marker to prove the source reaches the expected context.
3. Change one field, encoding, parser, or carrier at a time using a context-specific safe probe.
4. Distinguish filter/WAF behavior from application execution.
5. Follow the result to the final proof surface: database-derived value, command output, callback, safe file canary, browser runtime effect, persisted readback, internal-service response, or served upload behavior.
6. Repeat with a negative control and preserve exact request/action and response/proof artifacts.

If runtime proof is unavailable, report the full reachable code path and missing precondition as a candidate or code-level risk. Do not claim confirmed exploitability.

## Apply class-specific proof gates

Read [exploitability-proof-gates.md](references/exploitability-proof-gates.md). Enforce the appropriate gate before confirming a finding. In particular:

- require database-derived extraction for SQL injection;
- require browser/runtime JavaScript execution for XSS;
- require exact-sink callback, internal response, metadata, or file/protocol proof for SSRF;
- require command output or controlled callback for command/RCE claims;
- separate upload acceptance from browser, server, or document-processing exploitation.

Keep timing, errors, reflection, status/size differences, accepted files, listener presence, permissive headers, scanner labels, and source-only hypotheses in a candidate ledger until the gate passes.

## Review the control at the right layer

For each path, determine whether the defense is structurally correct:

- use parameterized APIs instead of blacklist filtering;
- validate canonicalized data at the authoritative server boundary;
- allowlist URL schemes/hosts and revalidate every redirect and resolved address;
- disable unnecessary parser features and unsafe polymorphic deserialization;
- isolate file storage and processing, generate server-side names, and serve inertly;
- use context-specific output encoding and safe DOM APIs;
- enforce origin and message schema checks before acting on cross-window data.

Search for variants of the same unsafe pattern across the repository after proving one path.

## Report proof and coverage

For every confirmed issue, provide source-to-sink path, exact entrypoint, control failure, safe proof, impact, affected variants, remediation location, and reproduction steps. Redact secrets and PII while retaining field names, types, lengths, fingerprints, hashes, and proof signals.

List untraced sources, unexecuted sinks, parser variants, background jobs, browser-only paths, unavailable callbacks, missing disposable fixtures, and other coverage gaps. Do not convert incomplete tracing into a clean result.

Referenced files: 3

tahr-verify-security-fix5.79 KB

View saved version →

---
name: tahr-verify-security-fix
description: Retest a security fix in the exact vulnerable context, decide whether the exploit path is closed, and validate secure remediation and regression coverage without breaking legitimate behavior. Use after a vulnerability patch, remediation commit, PR fix, dependency or configuration change, failed security retest, or when developers need proof that a fix is complete rather than a superficial code change.
---

# Tahr Verify Security Fix

Verify the failed security control and the attacker outcome, not merely the
presence of a patch. Preserve the original finding, proof, context, and scope so
a different identity, tenant, route, payload, deployment, or render path cannot
produce a false pass.

## Establish retest authority and safety

1. Record the original finding ID, exact claim, required proof, original
   positive evidence, affected revision, fix revision, and supplied patch.
2. Record the original actor/role, owner/tenant/object, endpoint or workflow,
   method/content type, payload class, state, configuration, render/trigger
   context, and final impact.
3. Default to source and test inspection. Exercise only an explicitly
   authorized local or staging target. Do not infer permission for production,
   third-party, destructive, credential-changing, billing, or broad data tests.
4. Use disposable accounts and fixtures. Preserve protected identities and
   extract only the minimum proof sample.
5. Redact passwords, cookies, bearer/session tokens, API keys, private keys,
   reset codes, personal data, and customer data. Retain only type, location,
   and hash/fingerprint when necessary.

Read [exact-context-retest-verdicts.md](references/exact-context-retest-verdicts.md)
and create separate candidate, proof, and coverage records before testing.

## Inspect the remediation

Trace the original source-to-sink or missing-control path through current code.
Identify:

- the exact failed control and where the patch now enforces it;
- existing project policy, ownership/tenant scope, validator, sanitizer,
  encoder, parameter binding, allowlist, network guard, or safe API reused;
- adjacent route, resolver, service, repository, serializer, worker, webhook,
  content type, method, redirect, or render path that may bypass the fix;
- public behavior, auth semantics, tenant rules, response shape, data
  invariants, and framework lifecycle that must remain compatible.

Search for a concrete contradiction to the fix claim. A new helper, regex, test,
or guard name is not proof that the effective path uses it.

## Retest in the exact context

1. Reproduce the original benign baseline.
2. Replay the original exploit or closest safe equivalent using the original
   actor, ownership/tenant relation, route, content type, state, and trigger.
3. Confirm the expected denial, encoding, validation, scoped result, blocked
   network/file action, safe query, or non-execution result.
4. Read back durable state or observe the final sink. Do not stop at the first
   accepting or rejecting layer.
5. Run a positive control proving the legitimate owner, tenant, role, input, or
   workflow still succeeds.
6. Run class-relevant alternate encodings, methods, content types, object IDs,
   roles, tenant relationships, render contexts, redirects, or concurrency
   variants within the safe budget.
7. Test sibling paths that share the changed control.

Use fresh, identity-matched sessions and prove object/tenant ownership for
authorization retests. A `200`, body-size change, reflection, upload acceptance,
tool success flag, timeout, stale session, or WAF response is not a retest
verdict by itself.

## Assign an exact verdict

Assign one verdict from the reference contract:

- `FIX_VERIFIED`;
- `FIX_PARTIAL`;
- `NOT_FIXED`;
- `REGRESSION_INTRODUCED`;
- `INCONCLUSIVE`.

Support it with finding-local positive evidence, negative evidence, controls
tested, contradiction result, exact-context match, variants, limitations, and
final basis. Tooling or environment failure alone requires `INCONCLUSIVE`, not
a pass or failure.

## Remediate securely when authorized

If the user asks only for verification, report the defect and do not edit. If
the user also asks to fix it, read
[remediation-and-regression-gates.md](references/remediation-and-regression-gates.md)
and:

1. Localize only real repository files, symbols, routes, controls, and tests.
2. Repair the failed control at the smallest correct shared layer.
3. Reuse established project security primitives before creating parallel
   logic.
4. Avoid blacklist/regex-only fixes when parameterized, encoded, scoped,
   schema-validated, or allowlisted APIs exist.
5. Preserve public behavior except for the intentional rejection of the
   vulnerable action.
6. Add a regression test that fails on the vulnerable behavior and a positive
   control for legitimate behavior.
7. Run focused tests first, then feasible lint, type, build, integration, and
   broader tests.
8. Review the final diff for adjacent bypasses, superficial controls, generated
   edits, unrelated churn, secret leakage, and compatibility breaks.

Do not mark the fix ready because a test was added; confirm the test reaches
the original sink or protected action and fails without the effective control.

## Account for coverage and conclude

Account for the original path, every material alternate path, sibling consumer,
security-relevant variant, required negative control, legitimate positive
control, and test command. Give unresolved high-risk rows an exact blocker and
owner.

Do not issue `FIX_VERIFIED`, “ready,” or a clean conclusion while an original
proof element is missing, the exact context was not reproduced, a high-risk
bypass path is unreviewed, required tests did not run, or a material regression
remains. Use `FIX_PARTIAL`, `REGRESSION_INTRODUCED`, or `INCONCLUSIVE` and state
the next evidence needed.

Referenced files: 3

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 12:00 UTC
Collection status
Collected

plugins_6aab2358993881919183acd471020907

Download listing JSON