← Plugin catalog
Security
Tahr Security
Tahr Security Inc v0.3.3
Review applications with an evidence-first security workflow that maps attack surfaces, tests authentication and access controls, traces dangerous inputs, models threats, and verifies fixes.
Language: English · Automatically detected from descriptions.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- GPL-3.0-only
- Package author
- Yack Security Inc
Declared capabilities
- Read
- Write
Package observed Sep 30, 2026.
Files & skills
File archives
Plugin package81 files · 347 KBBrowse files →
Skill instructions
tahr-audit-android5.36 KB
--- name: tahr-audit-android description: Audit Android application security from an APK, AAB-derived APK, Android source repository, manifest, or authorized emulator/device. Use for mobile release reviews, OWASP MASVS-oriented assessments, exported component and deep-link testing, WebView and IPC review, local storage and token analysis, mobile API traffic review, runtime instrumentation, privacy testing, and Android hardening validation. --- # Tahr Audit Android Use an artifact-first workflow to turn manifest and code signals into focused runtime checks. Keep static evidence, runtime reachability, and verified security impact distinct. ## Set scope and safety 1. Identify the application ID, build variant, APK hash, source revision, device profile, identities, and backend environment in scope. 2. Treat a local repository or supplied artifact as authorized for read-only review. Default to static analysis when active device or backend testing is not clearly authorized. 3. Use a disposable emulator snapshot, test install, test accounts, and synthetic data for state-changing checks. 4. Do not clear package data, change protected account credentials, trigger lockouts, submit real payments, send messages, write through content providers, load persistent code, or tamper with production data. 5. Do not modify the app or implement fixes unless the user explicitly asks. Use secrets, tokens, PII, keys, cookies, and credentials only transiently for authorized verification. Persist their class, source, key or field name, redacted excerpt, length, fingerprint, scope, expiry, and replay result—not raw values. ## Inventory before testing Read [attack-surface.md](references/attack-surface.md) before exploring the application. Prefer existing artifacts over broad rescans: manifest, decompiled source, smali/resources, class or string index, static findings, device information, traffic capture, dynamic plan, and earlier coverage. If no compact index exists, create a focused inventory of URLs, secrets, crypto APIs, log calls, password/token fields, WebViews, JavaScript bridges, dynamic loading, and native libraries. Build a target list with provenance for: - exported activities, services, receivers, providers, and their permissions; - deep links, app links, intent actions, URI authorities, and parameters; - WebViews, loaded origins, settings, file/content access, and bridges; - authentication, biometric, token, OAuth, and session paths; - storage locations, backup behavior, logs, clipboard, and caches; - crypto operations, keys, IVs, RNG, signing, and integrity decisions; - endpoints, trust configuration, cleartext paths, pinning, and PII flows; - deserialization, reflection, commands, dynamic code, native libraries, and dependencies; - privacy permissions, tracking SDKs, consent states, retention, and screenshots. ## Separate evidence levels Classify each observation immediately: 1. **Static candidate** — source, smali, manifest, resource, dependency, or configuration evidence identifies a plausible weakness. 2. **Runtime reachability** — adb, UI, proxy, logcat, filesystem, Frida, or backend evidence shows that the path executes or is externally callable. 3. **Verified impact** — the behavior crosses a confidentiality, integrity, authorization, authentication, privacy, or protected-functionality boundary. Never label a static flag as runtime exploitation. An exported component, weak algorithm, permissive WebView setting, cleartext allowance, dependency CVE, or disabled resilience control is a target until the matching proof gate is met. ## Plan and run focused checks Read [proof-gates.md](references/proof-gates.md) before dynamic testing or severity assignment. Prioritize high-impact targets from artifacts. For each target, define the safe command or interaction, caller and app state, expected secure behavior, impact proof, cleanup, and stop condition. Exercise fresh, logged-in, proxied, unproxied, and offline states only when they add relevant evidence. Use read-only or marker-based component and provider checks by default. Instrument only the classes and methods identified by static evidence. Capture metadata rather than raw values. Stop when the device, app, backend, or account shows instability. For dependency or platform intelligence, search only exact detected packages and versions. Record affected range, fixed version, prerequisites, and source, then prove that the vulnerable code path is packaged and reachable. Intelligence is a lead, never a finding. When stuck, use a bounded fresh-agent pass to suggest missing test states or alternate proof paths. Treat suggestions as candidates and independently verify them. ## Report without losing uncertainty Return three separate sections: - **Verified findings:** include category proof, exact component or file, caller/app state, commands or interactions, before/after behavior, impact, and remediation. - **Candidates:** include the static or runtime signal, missing proof, and safest next check. - **Coverage:** mark every material target `tested`, `partially_tested`, `skipped_with_reason`, `inaccessible`, `blocked_by_environment`, or `unsafe_or_destructive_skip`. Do not silently omit high-risk components because the emulator, login, proxy, root, Frida, or backend was unavailable. Do not interpret a blocker as evidence that the app is secure. Calibrate severity to demonstrated impact, not MASVS category names or scanner labels.
Referenced files: 3
tahr-audit-secrets-config7.05 KB
--- name: tahr-audit-secrets-config description: Audit application-owned secrets, cryptography, dependency reachability, infrastructure-as-code, containers, CI/CD, cloud permissions, and runtime security configuration with evidence and false-positive controls. Use for repository hardening, deployment review, leaked-key triage, dependency/CVE review, exposed debug or admin surface checks, or pre-release configuration audits. --- # Tahr Audit Secrets and Config Find configuration and supply-chain weaknesses that become real attacker capabilities. Do not turn a keyword, permissive development setting, or package advisory into a vulnerability without proving production relevance and reachability. ## Establish scope and safety Inspect source, committed configuration, lockfiles, build files, IaC, containers, CI/CD, and documentation read-only by default. Inspect git history only when the user includes it. Do not search unrelated home directories, credential stores, or `.git` working metadata. Never validate a discovered credential against a live provider unless the user explicitly authorizes that exact action and the account is disposable. Never print, copy, commit, or persist a raw secret. Record type, file and line, source, scope, length, a short redacted preview, and a SHA-256 fingerprint. Read [references/config-audit-matrix.md](references/config-audit-matrix.md) for the audit families. Read [references/exploitability-research.md](references/exploitability-research.md) before using advisory or CVE information. ## Inventory every configuration plane Enumerate: - application config and environment loading, including defaults and fallbacks; - secret-manager, KMS, key-vault, certificate, and signing-key integrations; - manifests and lockfiles for production, development, plugins, images, actions, and build tools; - Dockerfiles, compose files, Kubernetes, Helm, Terraform, Pulumi, CloudFormation, Ansible, systemd, IIS, reverse proxies, and serverless definitions; - CI workflows, release jobs, artifact publishing, package registries, caches, and deployment scripts; - browser/mobile public configuration, source maps, runtime config endpoints, health, metrics, debug, admin, docs, and actuator surfaces; - logging, telemetry, backups, data exports, error handling, and crash artifacts. Mark generated, vendored, example, fixture, test-only, local-only, and production-relevant paths. Missing deployment material is a coverage gap, not proof of secure deployment. ## Review secret exposure Search provider-specific formats and contextual assignments for cloud keys, OAuth clients, signing/session/webhook secrets, database and queue URLs, private keys, developer tokens, payment keys, AI provider keys, backup credentials, and privileged API keys. For each hit: 1. Determine whether it is application-owned committed content, committed history, generated output, a placeholder, documentation, test fixture, environment reference, secret-manager lookup, public browser key, or local checkout metadata. 2. Exclude `.git/config`, remote URLs, hooks, logs, and scanner checkout credentials from application findings. A secret genuinely committed in repository history remains in scope. 3. Determine the exposure path: shipped client bundle, public artifact, image layer, CI log, repository audience, runtime endpoint, backup, or developer-only file. 4. Determine likely privilege, environment, restrictions, rotation status, and blast radius without using the raw value. 5. Inspect whether the application fails closed when the secret is absent or falls back to a default, empty, or hardcoded value. 6. Classify the result as `verified`, `static-confirmed`, `candidate`, or `rejected` with exact reasons. Do not report environment-variable references or vault lookups as hardcoded secrets. Treat public client identifiers and publishable keys according to provider design; require missing restrictions or privileged use before claiming impact. ## Review production controls Trace configuration from source default through environment override to deployed consumer. Look for: - debug/test modes, stack traces, verbose errors, install/setup routes, sample accounts, default credentials, and unrestricted metrics or admin endpoints; - fail-open authentication, authorization, origin, webhook-signature, feature-flag, or network-policy behavior when configuration is missing or malformed; - wildcard or reflected credentialed CORS, unsafe cookie/session flags, proxy trust, generated-link host trust, weak TLS, missing transport enforcement, and cache-key confusion; - overly broad IAM, public storage, unauthenticated services, exposed databases, unrestricted egress, cloud metadata access, and security-group/network-policy gaps; - privileged containers, root users, dangerous capabilities, host namespaces, writable mounts, Docker socket exposure, unpinned images, and secrets embedded in layers; - untrusted pull-request code reaching privileged CI secrets, mutable third-party actions, artifact poisoning, unsafe interpolation, and excessive workflow permissions; - sensitive logging, personal data in URLs, backups without access controls, client-side secret storage, and long retention; - weak password hashing, encryption, randomness, key/nonce/IV reuse, insecure verification, or custom cryptography tied to a real security property. Read the code or deployment path that consumes each setting. A permissive example file or a development-only branch is not a production finding unless it can affect a shipped environment. ## Establish dependency reachability For every dependency lead: 1. Prove the exact resolved version from a lockfile, image digest, installed artifact, or reproducible build evidence. 2. Identify the affected function, class, feature, plugin, image component, action, or transitive path. 3. Prove first-party use and a reachable entrypoint, job, parser, upload, request, build, or deployment path. 4. State the attacker-controlled input or prerequisite. 5. Inspect wrappers, disabled features, sandboxing, network controls, version backports, and vendor fixes. 6. Separate `version-candidate`, `reachable-candidate`, `attempted`, and `proven` states. Do not report a CVE because a package name appears. If runtime reachability cannot be established, produce a precise upgrade/hygiene note or validation target rather than an exploitable finding. ## Challenge and report For every accepted claim, name the exact failed control and seek the strongest contradiction. Distinguish severity from confidence. Link configuration primitives into an attack chain only when every hop is independently proven. Report: - evidence-backed findings with redacted proof, production relevance, affected capability, remediation target, and regression or policy test; - candidates needing runtime, cloud, build, or owner confirmation; - rejected hits with placeholder/test/dev/framework control evidence; - coverage by configuration plane, including missing deployment artifacts and unreviewed environments. Do not say an application or deployment is secure when production configuration, runtime identity, cloud state, or dependency reachability was unavailable.
Referenced files: 3
tahr-map-attack-surface4.77 KB
--- name: tahr-map-attack-surface description: Map the real security-relevant surface of a web application or API from source, specifications, JavaScript, browser behavior, and authorized traffic. Use for pre-pentest reconnaissance, security-review scoping, hidden route or parameter discovery, undocumented API inventory, role-aware surface comparison, or judging whether an existing review actually covered the application. --- # Map the Attack Surface Build an evidence-backed inventory before testing vulnerabilities. Treat every discovered item as coverage evidence, not as a finding. ## Set the boundary 1. Identify the repository roots, application origins, API origins, environments, and supplied specifications or traffic captures. 2. Record the allowed runtime scope. Do not contact a live target unless the user supplied it or clearly authorized testing it. 3. Default to source-only analysis when authorization, credentials, or a runnable environment are absent. 4. Keep runtime activity read-only and low volume. Do not submit destructive forms, create real orders, send invitations, modify accounts, or enumerate unrelated infrastructure. 5. Never print or persist passwords, cookies, bearer tokens, API keys, reset links, private keys, CSRF values, or user PII. Retain names, locations, value classes, lengths, and short SHA-256 fingerprints when useful. ## Inventory independent evidence sources Inspect each available source independently before merging: - server routes, controllers, RPC handlers, middleware, authorization declarations, background jobs, queues, and WebSocket/SSE handlers; - OpenAPI, Swagger, GraphQL schemas and documents, generated clients, protobufs, and API examples; - frontend routes, forms, fetch/axios clients, lazy chunks, source maps, feature flags, upload configuration, storage keys, and custom headers; - infrastructure routes, reverse-proxy rules, serverless functions, storage buckets, callback handlers, and public documentation; - authorized unauthenticated and authenticated browser/API traffic, separated by identity, role, and tenant. Do not infer reachability from a route name alone. Mark each operation as source-derived, documented, browser-observed, traffic-observed, or runtime-confirmed. ## Build canonical operations Read [surface-inventory.md](references/surface-inventory.md) and create one record per meaningful operation. Preserve: - protocol, origin, method or operation type, path, content type, and request shape; - query, path, body, header, cookie, form, multipart, and GraphQL variable fields; - authentication state, identity/role/tenant context, object identifiers, owner hints, and workflow state; - response class, state-changing behavior, source artifact, and confidence. Do not merge operations when method, body shape, parser, auth state, role, tenant, owner, response behavior, or workflow state differs. Preserve exact endpoint-to-field provenance; a global parameter list is insufficient. ## Prioritize attacker-relevant surface Rank concrete operations higher when they expose: - login, reset, MFA, invite, token, session, or account lifecycle behavior; - admin, role, tenant, membership, ownership, billing, entitlement, approval, or settings functions; - object IDs in any carrier, bulk/composite IDs, exports, downloads, or sensitive response fields; - upload, import, preview, render, conversion, callback, webhook, URL-fetch, or integration behavior; - price, total, quantity, discount, status, state, owner, tenant, role, or idempotency fields; - HTML/Markdown/template input, search/filter/sort expressions, file paths, XML, or serialized data; - undocumented versions, hidden client routes, debug endpoints, source-map leads, or role-specific discrepancies. Priority controls review order only. Do not discard lower-ranked operations. ## Close coverage gaps Read [coverage-gates.md](references/coverage-gates.md). Assign every discovered item one terminal state: `visited`, `runtime_confirmed`, `source_only`, `queued`, `partially_explored`, or `skipped_with_reason`. When discovery is unexpectedly shallow, perform one bounded second pass using a different evidence source: follow lazy routes, inspect request builders, parse schemas, expand safe UI elements, or compare another supplied identity. Never describe absent evidence as proof that a feature is absent. ## Deliver the inventory Report: 1. scope and evidence sources inspected; 2. canonical operations grouped by trust boundary and identity context; 3. high-value targets with exact provenance; 4. identity, tenant, object, upload, callback, and workflow maps; 5. coverage gaps, blocked areas, and the consequence of each gap; 6. candidate hypotheses clearly labeled as unverified. Do not assign vulnerability severity from recon alone. Recommend the appropriate Tahr testing skill for each candidate class.
Referenced files: 3
tahr-review-tahr-findings3.83 KB
---
name: tahr-review-tahr-findings
description: Read applications, assessments, and findings from an already configured Tahr MCP connection. Trigger only when the user explicitly asks to query, list, summarize, or review Tahr account data; do not trigger for generic security reviews, source-code reviews, or non-Tahr findings.
---
# Review Tahr Findings
Use this optional, read-only workflow for existing Tahr customers who have manually configured the Tahr MCP server in Codex. This is a Codex/local manual integration only. Never claim that it provides public ChatGPT account linking.
## Establish account context
Start every explicit Tahr-data request with `get_context`. Before querying account data, confirm and report the authenticated organization name and the relevant capabilities. If the user named a different organization, stop and ask them to switch or reconfigure the connection; do not query the authenticated organization.
Respect the reported capabilities:
- Access applications only when `canReadApplications` is true.
- Access findings only when `canReadFindings` is true.
- Treat this workflow as read-only even when `canEditFindings` is true.
Use only these tools: `get_context`, `list_applications`, `get_application`, `list_assessments`, `list_findings`, and `get_finding`. Never invoke mutation tools or perform write, triage, status, or comment actions.
## Handle connection failures safely
If the MCP server or tools are unavailable, authentication returns 401, the token is missing, revoked, or expired, permission fails, or the service is unavailable or rate limited, give concise setup or recovery guidance. Suggest checking that the personal token environment variable and Codex MCP configuration are present, restarting Codex, obtaining the required organization access, or retrying later as applicable. Never ask the user to paste a token into chat. Offer to continue with the independent local security skills.
## Resolve and query records
Resolve a human application name with `list_applications`, then use the exact application ID returned by Tahr. Do not reveal details for null, missing, or cross-organization resources.
List tools accept optional `cursor` and `limit` parameters. Use a limit from 1 through 50; the default is 25. Responses contain `items` and `page.{nextCursor,isDone}`. Paginate only when the user requests all or complete results. For each subsequent page, keep every filter unchanged and pass the prior `nextCursor`. Otherwise, stop when the request is satisfied and disclose the coverage limit.
For `list_findings`:
- Always provide `kind` as `security` or `authorization`.
- Use the default `assessmentScope: latest` unless the user explicitly requests history or all assessments; only then use `assessmentScope: all`.
- Filter as needed by `applicationId`, `assessmentId`, or severity: `Critical`, `High`, `Medium`, `Low`, or `Info`.
- When generic "findings" clearly means both security and authorization findings, query each kind separately. Otherwise, clarify the ambiguity before querying.
Use `get_finding` only when requested or needed for the requested detail. Treat returned records as Tahr platform records, not independently verified vulnerabilities. Preserve finding IDs and assessment IDs in summaries where helpful, and distinguish recorded evidence from claims.
If list items are empty, report only that no matching records were returned; do not speculate. If a get operation returns null, report that the resource is unavailable without disclosing whether it exists elsewhere.
## Protect data and report results
Do not send repository content, secrets, credentials, tokens, or unrelated chat data to Tahr. Never log or persist the bearer token.
Provide concise result summaries grouped by severity or application as requested. State whether the summary covers all matching pages or only the pages and item limit queried.
Referenced files: 1
tahr-secure-app7.76 KB
--- name: tahr-secure-app description: Perform an evidence-backed, pentester-style security review of an application from source, configuration, specifications, tests, and optionally an explicitly authorized local or staging runtime. Use for comprehensive app security audits, pentest readiness, pre-release reviews, dangerous-flaw discovery, or coordinating the Tahr specialist skills; also use when a prior scanner or LLM review created confidence that needs independent verification. --- # Tahr Secure App Review the application as an attacker would: map what is reachable, form target-specific hypotheses, try to disprove each candidate, require exploit-class proof, and account for what was not tested. ## Set the review mode Choose one mode and state it before reviewing: - `deep`: inspect every admitted first-party file and close every high-risk gap. Use by default for “secure this app.” - `focused`: review a named feature or boundary deeply and list excluded areas. - `retest`: verify an accepted fix with `$tahr-verify-security-fix`. Treat source and local artifact inspection as read-only review. Run runtime probes only against a local, disposable, or explicitly authorized test target. Do not infer permission to test production, third parties, other tenants, or real accounts. Use low-impact canaries, disposable objects, bounded concurrency, and reversible actions. Never persist raw passwords, tokens, cookies, private keys, personal data, or cloud credentials. ## Create the evidence ledger Read [references/evidence-contract.md](references/evidence-contract.md) before classifying any issue. Use [assets/review-report.template.json](assets/review-report.template.json) when a durable JSON artifact is useful. Write artifacts only where the user requests or in a clearly named local review directory; do not modify application code during a review. Keep three lanes separate throughout the work: 1. `findings`: claims that meet the applicable proof status. 2. `candidates`: concrete leads whose required proof is incomplete. 3. `coverage`: reviewed, partial, blocked, skipped, and not-applicable scope. Never promote a scanner hit, dangerous function name, package advisory, route name, response status, reflection, timing change, accepted upload, or prior report statement by itself. ## Map before judging Inventory the application before searching for bugs: - first-party files, frameworks, manifests, lockfiles, infrastructure, deployment and CI configuration; - HTTP routes, GraphQL resolvers, RPC/gRPC handlers, WebSockets/SSE, webhooks, uploads, callbacks, serverless functions, CLI entrypoints, queues, jobs, and externally influenced schedulers; - actors, roles, service principals, tenants, organizations, groups, ownership fields, sessions, and recovery flows; - assets, data stores, secrets, billing or entitlement state, admin/debug functions, AI models, RAG sources, tools, and external integrations; - browser routes, lazy chunks, source maps, API clients, mobile deep links, and undocumented or legacy API versions. Follow thin controllers into middleware, policies, services, repositories, serializers, templates, and asynchronous consumers. Mark each admitted item `reviewed`, `partial`, `blocked`, `skipped-with-reason`, or `not-applicable`. Low apparent risk is a reason to review briefly, not to disappear the file from coverage. Use `$tahr-map-attack-surface` when the reachable surface or identity coverage is unclear. ## Build attacker hypotheses Create target-specific hypotheses in four forms: - `actor -> action -> resource -> owner/tenant boundary -> expected decision`; - `untrusted source -> transformations -> dangerous sink -> expected control`; - `workflow state -> attempted transition/replay/race -> invariant -> authoritative readback`; - `deployment input/default -> privileged capability or sensitive asset -> compensating control`. Prioritize unauthenticated paths, cross-user or cross-tenant boundaries, state-changing functions, sensitive exports, callbacks and outbound fetches, file processing, server-side rendering, privileged fields, legacy versions, recovery paths, background workers, and AI tool use. ## Route to specialist skills Read [references/review-routing.md](references/review-routing.md) and invoke only the applicable specialist skills. A comprehensive review normally includes: - `$tahr-test-authentication` for login, recovery, MFA, OAuth/OIDC, tokens, cookies, and session lifecycle; - `$tahr-test-access-control` for a complete actor-resource-action model, path-specific authorization traces, proof-gated findings, and safe validation; - `$tahr-trace-dangerous-inputs` for injection, browser sinks, outbound requests, parsers, files, and uploads; - `$tahr-test-business-workflows` for state, pricing, quota, invitation, approval, replay, and race abuse; - `$tahr-audit-secrets-config` for secrets, cryptography, dependencies, infrastructure, and deployment defaults; - `$tahr-test-ai-agents` or `$tahr-audit-android` when those technologies exist; - `$tahr-threat-model-app` when a full-system threat model and validation plan are requested for the entire existing application. ## Investigate and challenge For every candidate: 1. Cite the exact route, file, line or symbol, actor, input, object, or configuration involved. 2. Trace reachability across files and processes; distinguish dead, test, generated, dependency, and production code. 3. Name the expected control and inspect its actual placement and applicability. 4. Search for the strongest contradiction: middleware, policy, tenant filter, validator, encoder, allowlist, safe parser, parameter binding, environment guard, or framework behavior. 5. Record the contradiction verdict as `not-contradicted`, `contradicted`, `partially-contradicted`, or `insufficient-evidence`. 6. Define the exact claim and proof needed before attempting runtime validation. 7. Establish a normal baseline and an expected-denial or benign negative control. 8. Use the smallest safe proof. Record request/state/browser/callback/readback evidence without raw secrets. 9. Search one bounded set of sibling routes, helpers, models, and variants after a strong seed; give each variant its own evidence. Operational failure is a limitation, not target-side proof. Stale authentication, missing roles, WAF or rate limiting, callback outage, tool error, or target instability must not become either a finding or a false-positive conclusion. ## Close coverage and chain impact Before concluding: - resolve every high-risk candidate as accepted, rejected with exact control evidence, out of scope, or deferred with a specific blocker; - list unreviewed files, endpoints, identities, tenants, workflows, sinks, environments, and runtime-only claims; - distinguish “tested with no issue observed” from “not tested”; - link only independently proven primitives into attack chains, and require evidence for every hop; - avoid “secure” or “no vulnerabilities” language when material gaps remain. A clean result means only: no verified findings were produced within the stated, completed scope. It is not a guarantee about omitted or blocked scope. ## Deliver the review Lead with dangerous, reachable issues. For each finding include the exact claim, affected location, preconditions, required proof, positive and negative evidence, contradiction result, impact, confidence, limitations, remediation target, and a repo-native regression test. Keep severity separate from confidence. Summarize candidates and coverage gaps after findings. Redact sensitive substrings while preserving type, source, length, scope, expiry where relevant, and a non-secret fingerprint. Validate a JSON review artifact with: ```bash python3 <skill-directory>/scripts/validate_review.py path/to/tahr-review.json ``` Use `--strict` only when claiming a deep review has no unresolved high-risk scope.
Referenced files: 5
tahr-test-access-control13.4 KB
--- name: tahr-test-access-control description: Perform complete or focused, evidence-backed access-control review from source and optionally an explicitly authorized local or staging runtime. Model subjects, roles, tenants, resources, actions, properties, policy rules, enforcement points, and owner-attributed test cases; trace object-, function-, property-, role-, and tenant-level authorization through REST, GraphQL, web, job, and asynchronous paths; safely validate IDOR/BOLA/BFLA, mass assignment, privilege escalation, and cross-tenant isolation; and reject status-code or guessed-ID false positives. Use for authorization code review, multi-user or multi-tenant assessments, admin and role boundary analysis, pre-pentest review, or validation of a suspected access-control finding. --- # Tahr Test Access Control Determine exactly who can perform which action on which resource, property, or function—and prove when the implementation violates that application-specific rule. Produce an authorization model and proof ledger, not a status-code diff. ## Load the operating contract Before reviewing: 1. Read [full-review-workflow.md](references/full-review-workflow.md) for scope, execution order, completion, and lifecycle rules. 2. Read [access-control-data-contract.md](references/access-control-data-contract.md) for stable IDs, exact enums, and record relationships. 3. Read [access-matrix.md](references/access-matrix.md) before constructing operation, relationship, carrier, property, or variant coverage. 4. Read [access-proof-gates.md](references/access-proof-gates.md) before accepting or rejecting a candidate. 5. Read [runtime-test-safety.md](references/runtime-test-safety.md) before any runtime action. 6. Read [source-review-patterns.md](references/source-review-patterns.md) when source is available. 7. Use [worked-example.md](references/worked-example.md) only when the expected evidence-to-finding trace is unclear. Start from [access-control-review.template.json](assets/access-control-review.template.json) and keep [access-control-review.schema.json](assets/access-control-review.schema.json) as the canonical output contract. Maintain one `access-control-review.json`; derive all reader-facing artifacts from it. ## Choose scope and assurance honestly Set `review_mode` to: - `full` for every admitted authorization-relevant operation in the existing application; or - `focused` for explicitly named operations, findings, resources, or policy boundaries. Default to `full` when the user asks to review or secure the application and does not explicitly narrow the authorization scope. Set `analysis_basis` to `source_only`, `runtime_only`, or `hybrid`. A focused review must carry a visible limitation and must not make an application-wide claim. A source-only review may be `complete` with `source_observed` assurance and planned runtime tests when every declared source surface is dispositioned. It must not claim that an attack ran or that deployed enforcement failed. Keep these states separate: - `review_status`: whether the declared authorization scope was dispositioned; - `assurance_status`: source observation versus authorized runtime validation; - candidate `disposition`: lead, follow-up, rejected, source-confirmed, or runtime-confirmed; - test `execution_status`: planned, passed, failed, inconclusive, or blocked; - coverage `status`: reviewed, tested, pending, deferred, or out of scope. Default to read-only analysis. Never start an application, send a request, refresh a session, or mutate state unless the exact target and action class are authorized. Never commit, patch, or reconfigure the reviewed application as part of this skill. ## Freeze the review inputs For source or hybrid review, create a deterministic manifest in the selected output directory: ```bash python3 <skill-directory>/scripts/build_review_manifest.py \ path/to/application --include . \ --output path/to/output/repository-manifest.json ``` Pass `--revision` for an immutable VCS or release revision. When omitted, the script derives `snapshot-sha256:<digest>` from admitted paths and bytes. Copy its embedded revision and content hash into metadata, manifest evidence, and coverage inventory. Repeat `--package` for package/module labels and `--document` for supplied repository policy or design files; documents are admitted and hashed automatically. Use explicit empty arrays for absent documents or exclusions. When the output lives under the application root, place it in a dedicated subdirectory such as `.tahr-review/`; the builder records and excludes that whole directory to prevent generated artifacts from contaminating later snapshots, and refuses output directly in the root. Inventory every admitted REST route, GraphQL query/mutation/subscription, server action, RPC method, UI-backed function, webhook, worker/job, queue consumer, export/download, bulk operation, legacy/versioned interface, and administrative surface that makes or depends on an authorization decision. Record intentionally public and non-applicable surfaces instead of deleting them from the inventory. When `$tahr-map-attack-surface` is available, consume its frozen operation inventory and reconcile it; otherwise inventory locally. Companion skills are optional—the access-control skill must remain independently usable. ## Establish identity and object truth Model unauthenticated, user, peer, role, tenant, administrator, support, service, worker, integration, and other applicable subjects. Keep source-modeled identities separate from runtime-verified sessions. For runtime identities, bind the redacted auth artifact fingerprint to an observed caller, role, tenant or authorization domain, transport, freshness, and validation evidence. A filename, configured label, JWT claim, profile field, or successful HTTP response alone does not make a session trustworthy. Mark unusable or mismatched identities as coverage limitations; never turn their failures into target-side denials. For each target object, record resource type, exact identifier carrier, owner or controller, tenant/domain, sensitivity, lifecycle, and provenance. A shared ID pool, guessed adjacent ID, globally discovered identifier, or caller-owned `/me` object cannot prove peer or cross-tenant impact. Use owner-attributed source evidence, an authoritative owner baseline, or a disposable object created and read back under the owner identity. Treat supplied assessment identities as protected. Do not delete, disable, lock, re-role, rename, reset, or rotate them. Use fresh disposable objects and non-protected accounts for state-changing validation. ## Build the authorization model before judging Record subject, action, resource, property, relationship, tenant/domain, workflow state, feature/plan, authentication strength, and other policy context. Express each rule as `allow`, `deny`, `conditional`, or `unknown` and cite its authority. Prefer source policy, middleware, domain rules, role matrices, documented requirements, or explicit assessment context. UI hiding, endpoint names, generic assumptions about administrators, and an owner-success baseline may corroborate a rule but do not normally prove expected denial alone. For every operation, enumerate each independent authorization obligation: source object, destination object, parent, child, relationship object, property, function, and asynchronous continuation. Record every caller-supplied identifier and sensitive property, where restrictions enter the path, where they are consumed, the authoritative query or state change, and every downstream enforcement point checked. ## Trace controls end to end When source is available, trace shipped entrypoint to final data return, mutation, worker, integration, signed URL, audit record, notification, or other side effect. A route-level guard that permits an action somewhere is not proof that the submitted target is in scope. A policy helper that exists but is not consumed on this path is not a control. Search for the strongest contradiction before retaining a gap: global middleware, dependency injection, decorators, domain policy, repository filters, ORM scopes, serializers, workers, database policy, deployment controls, and sibling route variants. Mark dead, test-only, generated, dependency, or unreachable paths explicitly. Group operations only when policy, resource, action, enforcement point, and relevant carriers and variants genuinely match. Preserve operation-level IDs so grouping cannot hide a legacy route, bulk path, alternate parser, nested GraphQL resolver, or asynchronous continuation. ## Construct complete matrix coverage For each applicable operation, freeze the expected callers, relationships, identifier/property carriers, and interface variants before recording results. Include unauthenticated, own-object, same-role peer, cross-role, cross-tenant, service-to-user, and privileged-function cases when the model makes them meaningful. Pair unauthorized cases with an authorized baseline of the same operation shape. Change one declared authorization dimension at a time. If an identity, owner-attributed object, parser, route variant, or safe fixture is unavailable, keep the matrix cell and mark the exact blocker; never omit it to improve coverage. Ranking controls execution order, not inventory. A full review processes every expected case or gives a specific disposition. Planned runtime validation does not by itself make completed source coverage incomplete; missing high-risk source analysis or an explicitly required runtime case does. ## Execute only safe, discriminating tests Follow [runtime-test-safety.md](references/runtime-test-safety.md). Require the test to match exactly one structured authorization target across origin, environment, surface/operation/resource IDs, tenant or domain, identities, transport, action and mutation scope, request/attempt limits, and validity window. Use synthetic data and the smallest reversible proof. Every test defines and distinguishes: - caller-identity proof; - owner/tenant/target attribution; - authorized baseline success; - expected denial or control-held signal; - unauthorized protected-data/action impact or control-failure signal; - authoritative readback and cleanup for state changes. An executed result uses one `run_id`. Its freshly collected identity preflights and every observed signal must cite same-run, same-target runtime evidence. Keep purpose-specific evidence for each signal; one aggregate record cannot be the sole proof for caller, target, baseline, denial/impact, readback, and cleanup. `planned` and `blocked` tests never contain a result. HTTP 200, non-empty output, size/hash differences, empty/null/false output, generic SPA shells, validation errors, 4xx/5xx reachability, 202 acceptance, or an echoed request are leads only. A mutation requires persistent authoritative readback or an equivalent side effect. Repeat an accepted runtime failure with fresh identity state and a fresh authorized control. ## Apply proof and false-positive gates Use the five mandatory gates: caller, target, ownership or tenant, expected denial, and unauthorized impact. A source-confirmed candidate must connect a shipped reachable entrypoint to a concrete protected data/action sink, show the missing or bypassed path-specific control, and record the strongest contradiction checked. A runtime-confirmed candidate must additionally cite a conclusive authorized test result with direct runtime evidence. Reject or retain as follow-up: - current-user endpoints returning only caller-owned or empty state; - legitimate sharing, public, support, or administrator behavior supported by an applicable policy; - guessed or shared IDs without owner attribution; - wrong, stale, mismatched, or transport-incompatible identity material; - soft 404s, shells, parser errors, resource absence, or unproven server errors; - writes without readback, async acceptance without a side effect, or results whose baseline changed more than the authorization dimension. Record rejected leads with the applicable control or contradiction evidence. Do not erase them; they demonstrate that the review challenged its own leads. ## Challenge and publish the model After the primary pass, use an independent subagent when available; otherwise perform a separate adversarial reasoning pass with fresh instructions. Require it to seek omitted operations and identities, unmodeled parent/child objects, unused restrictions, alternate routes/parsers/versions, false owner or tenant attribution, ambiguous collaboration, unsafe tests, unsupported impact, stale evidence, and hidden coverage gaps. Record every challenge and disposition. Validate the canonical model: ```bash python3 <skill-directory>/scripts/validate_access_control_review.py \ path/to/access-control-review.json --strict ``` Fix failures, rerun the challenger when material content changes, and render: ```bash python3 <skill-directory>/scripts/render_access_control_review.py \ path/to/access-control-review.json --output-dir path/to/output --strict ``` The renderer produces `access-control-review.md`, `validation-plan.json`, `findings.json`, and `coverage.json`. Do not edit derived artifacts as separate sources of truth. Lead with confirmed source or runtime findings, unresolved high-risk candidates, blocked tests, and coverage limitations. Never conclude that access control or the application is secure. A clean result means only that no additional proof-gated finding was produced within the declared completed scope and assurance level.
Referenced files: 13
tahr-test-ai-agents4.98 KB
--- name: tahr-test-ai-agents description: Test security boundaries in applications that use LLM chat, RAG or vector retrieval, memory, file or URL ingestion, model-rendered output, tool/function calling, MCP, or autonomous agents. Use for source-backed AI feature reviews, authorized local or staging runtime tests, prompt-injection assessments, cross-tenant retrieval checks, agent/tool abuse reviews, and AI resource-control testing. --- # Tahr Test AI Agents Review the application as an attacker crossing data, instruction, identity, rendering, and tool boundaries. Treat payload catalogs as aids; make the proof discipline the center of the assessment. ## Establish the boundary 1. Confirm the repository, feature, identities, and runtime target placed in scope. 2. Default to source-only analysis when authorization for active runtime testing is unclear. 3. Identify protected accounts, production data, third-party integrations, and actions that must remain read-only. 4. Do not modify application code, configuration, or deployed state unless the user separately requests remediation. 5. Use disposable tenants, documents, objects, tools, callback collectors, and marker values for active tests. Never delete data, transfer value, send real messages, publish content, rotate credentials, change privileges, exhaust a customer budget, or exfiltrate real private data. Bound concurrency, output length, request counts, and cost. ## Build the real attack surface Read [methodology.md](references/methodology.md) before selecting tests. Trace model and embedding SDK calls to their actual HTTP, WebSocket, queue, or background-job entry points. Include supporting routes for uploads, knowledge bases, conversations, memory, feedback, model settings, tools, and shared views. When source is incomplete, inspect OpenAPI, client bundles, forms, runtime requests, streaming frames, and error shapes. Confirm an endpoint only when it accepts AI-shaped input or produces generated, streaming, retrieval, embedding, model, or tool behavior. A route name containing `chat`, a generic JSON response, or a successful status code is not confirmation. For each confirmed surface, record: - endpoint, method, controllable field, and authenticated identity; - tenant, conversation, memory, and retrieval context; - model/provider and guardrail hints, marked `unknown` when unproven; - ingestion formats and retrieval sources; - renderer and downstream output sinks; - available tools, resources, prompts, and autonomous steps; - measured rate, token, output, streaming, and quota behavior; - evidence source and confidence. ## Test conditionally Establish a benign baseline before attack probes. Run only families whose preconditions exist: - test direct instruction override against an application-specific control; - test indirect injection only through content the application really ingests; - test retrieval and memory across validated user or tenant boundaries; - test output injection in the actual browser or downstream consumer; - test tool and MCP agency at real resource and authorization boundaries; - test resource controls with conservative measured probes; - attempt chains only from previously observed signals. Vary semantic attack concepts before cosmetic encodings. Record every attempt by endpoint, field, concept, payload hash, response class, and success criterion. For stochastic behavior, reproduce the same concept and boundary at least three times in five attempts unless the original security contract requires a stronger threshold. When blocked, ask at most a few fresh-agent passes for concise new concept axes or plausible chains. Treat that advice as leads only. Never let an adviser, model claim, or payload classification verify a finding. ## Apply proof gates Read [proof-gates.md](references/proof-gates.md) before promoting or scoring a result. Maintain three separate collections: 1. **Verified findings** — the required boundary and impact proof exists. 2. **Candidates** — a concrete signal exists, but state the missing proof and safest next check. 3. **Coverage** — list tested, partially tested, skipped, inaccessible, degraded, and unsafe-to-test surfaces with reasons. Do not promote model claims, marker-only obedience, generic policy text, stored-but-unretrieved content, API reflection, tool names, version intelligence, one lucky response, or theoretical cost. ## Handle evidence safely Use raw secrets or private data only transiently when authorized and necessary. Persist the value class, source, tenant or role context, redacted excerpt, length, fingerprint, endpoint, payload class, and reproduction count. Do not place raw tokens, cookies, system prompts, private documents, credentials, PII, or tool output into findings, screenshots, logs, or reports. Name findings after the demonstrated result, not the attempted technique. Calibrate severity to the proven data, action, resource, tenant, or browser boundary. End with prioritized remediation tied to the actual trust boundary and a concise residual-risk statement.
Referenced files: 3
tahr-test-authentication4.84 KB
--- name: tahr-test-authentication description: Review and safely test web authentication and session boundaries across login, registration, password reset, magic links, MFA or OTP, OAuth/OIDC, SAML, passkeys, tokens, cookies, logout, and recovery. Use for authentication code review, pre-release auth testing, account-takeover analysis, session-management review, SSO integration review, or validating an existing security assessment. --- # Test Authentication Find identity-boundary failures, not merely unusual responses. Treat auth material as untrusted until it proves the intended identity and transport. ## Establish safe scope 1. Identify source roots, runtime origins, supported authentication methods, supplied identities, and expected account lifecycle. 2. Do not send runtime requests unless the user supplied or authorized the target. Use source-only analysis otherwise. 3. Treat all supplied accounts as protected unless explicitly labeled disposable. Do not lock, reset, disable, delete, re-role, enroll or remove MFA/passkeys, rotate credentials, or invalidate all sessions on protected accounts. 4. Use invalid identifiers for low-volume response-shape checks and disposable accounts for lockout, reset completion, password changes, MFA mutation, code replay, and takeover proof. 5. Stop runtime testing on lockout text, CAPTCHA, rate limiting, disabled-account state, unexpected notification delivery, or unclear side effects. ## Prove identity truth first For each supplied identity: - observe the rendered login flow before submitting credentials; - classify password, split-step, OTP, MFA, magic-link, OAuth/OIDC/SAML, passkey, browser-bound, and custom stages; - identify hidden state, nonce, CSRF, tenant, organization, provider, or login-method choices; - validate success against an authenticated-only or identity-confirming endpoint; - record the observed user, role, tenant, auth mode, and whether cookies/tokens are portable or browser-bound. Do not equate a cookie, token, callback URL, HTTP 200, account picker, application shell, or pending MFA page with successful authentication. A failed role-specific login is a coverage blocker, not target access denial. ## Model the lifecycle Trace these state transitions when present: `registration/invite -> verification -> login -> step-up/MFA -> session refresh -> logout/revocation` `forgot-password -> delivery -> token/code validation -> password change -> prior-session behavior` `OAuth/SAML/passkey initiation -> provider/authenticator -> callback/completion -> application session` Record every endpoint, browser action, actor, token class, binding, one-time expectation, expiry, and alternate/mobile/legacy channel. Derive endpoints from source, specifications, JavaScript, and observed traffic before using fallback names. ## Execute the abuse matrix Read [auth-matrix.md](references/auth-matrix.md). Prioritize tests that cross an identity boundary: - valid-disposable versus invalid account enumeration controls; - rate limiting and weaker alternate endpoints; - pre-auth, post-password, post-MFA, and fully authenticated stage skipping; - token/code replay, wrong-account binding, stale-token reuse, and parallel requests; - recovery, factor, password, email, and security-setting changes without reauthentication; - session fixation, logout/timeout invalidation, CSRF, cookie scope, and refresh rotation; - OAuth redirect/state/nonce/PKCE/code/client/scope binding; - credentials or reusable secrets in URLs, responses, logs, JavaScript, caches, or browser storage. Change one dimension at a time and pair every abuse attempt with a valid control. Refresh or re-establish the exact identity after an intentionally invalidating test before interpreting later responses. ## Apply proof gates Read [auth-proof-gates.md](references/auth-proof-gates.md). Keep endpoint discovery, header observations, raw tokens, configuration smells, status codes, timing, and script labels in a candidate ledger until the class-specific proof gate passes. For every confirmed issue, preserve: - exact flow stage, endpoint/action, and actor/session context; - baseline and manipulated request or browser action; - token/cookie class and state using redacted fingerprints; - authenticated-only data/action, wrong-account binding, replay, persistent state change, or other concrete impact; - safe, faithful reproduction steps and cleanup or restoration status. Never persist raw passwords, cookies, bearer/refresh tokens, authorization codes, reset/magic links, OTPs, SAML assertions, passkey material, secrets, or PII. ## Report coverage honestly Separate `confirmed`, `candidate`, `not_reproduced`, `blocked_for_safety`, and `not_applicable`. List every untested lifecycle stage, missing disposable identity, browser-bound limitation, unavailable delivery channel, stale session, or provider blocker. Do not turn incomplete auth coverage into “no issue found.”
Referenced files: 3
tahr-test-business-workflows5.66 KB
--- name: tahr-test-business-workflows description: Model and safely abuse-test stateful business workflows, API operations, and application invariants such as checkout, billing, credits, invitations, approvals, entitlements, exports, uploads, integrations, quotas, and asynchronous jobs. Use for business-logic review, race-condition and replay testing, mass-assignment or excessive-property review, workflow bypass analysis, API version/parser comparison, or pre-pentest testing of critical product flows. --- # Test Business Workflows Look for normal-looking requests that produce outcomes the business never intended. Prove the outcome, not merely that an unusual value was accepted. ## Bound the work safely 1. Identify source roots, specifications, runtime scope, supplied identities, sandbox/test mode, and critical workflows. 2. Do not send live requests without clear authorization. Use source, tests, schemas, and captured traffic to model workflows otherwise. 3. Treat supplied accounts, customer data, billing objects, orders, balances, coupons, invitations, and integrations as protected. 4. Require a disposable fixture or explicit dry-run/preview/test mode before creating orders, changing plans or roles, transferring value, consuming benefits, sending messages, registering webhooks, or mutating persistent state. 5. Define before state, expected effect, authoritative readback, restoration, and cleanup before every mutation or race batch. Stop on unexpected side effects or instability. ## Reconstruct the real workflow For each critical workflow, combine code, UI, schemas, tests, background jobs, JavaScript, and authorized traffic. Record: - actors, roles, tenants, objects, and ownership; - entry conditions and server-side prerequisites; - states, transitions, terminal states, and asynchronous steps; - authoritative values and where each value originates; - one-time tokens, idempotency keys, approvals, expiries, counters, and quotas; - compensating actions, cancellation, rollback, and cleanup; - alternate API versions, clients, content types, bulk operations, and direct endpoints. Write explicit invariants such as “the server calculates the total,” “only the current owner may approve,” “a code is redeemed once,” or “a transition requires the immediately preceding state.” Cite the evidence for each invariant; do not invent product rules from route names. ## Derive an abuse plan Read [workflow-abuse-matrix.md](references/workflow-abuse-matrix.md). Cover applicable classes: - boundary values, type confusion, omitted fields, extra fields, nested/array forms, and duplicate parameters; - client-controlled price, total, discount, tax, quantity, balance, owner, tenant, role, status, approval, entitlement, or feature fields; - direct access to later steps, skipped/reordered prerequisites, repeated actions, stale token/state reuse, and alternate actors; - duplicate submission, idempotency collisions, last-byte/single-packet concurrency in a disposable environment, and time-of-check/time-of-use gaps; - rate/quota limits on sensitive successful operations; - version, method, parser/content-type, mobile/legacy, bulk, GraphQL alias/batch, and asynchronous variants; - excessive response properties, unsafe third-party responses, webhook signature/event handling, and generated-link/cache-key trust. Rank high-impact invariants first, but give every discovered workflow or operation a terminal disposition. Do not replace target-derived requests with guessed endpoints or generic bodies. ## Execute paired experiments For authorized runtime work: 1. Capture a valid single-request baseline and define what proves business success. 2. Change one invariant-related dimension while holding actor, object, method, body, parser, and timing constant. 3. Use a negative control and, where useful, an authorized control. 4. Read the authoritative state after the attempt; also inspect related list, audit, balance, entitlement, or downstream job state. 5. For races, first prove sequential behavior, then run the smallest bounded synchronized batch and count business successes—not merely HTTP responses. 6. Restore or delete only disposable fixtures and record cleanup. An accepted value, 2xx, redirect, response-size change, missing header, or multiple concurrent responses is candidate evidence until the unsafe business outcome is shown. ## Apply proof gates and chain impact Read [workflow-proof-gates.md](references/workflow-proof-gates.md). Confirm only reproducible outcomes such as unauthorized state transition, financial/value manipulation, duplicate benefit, quota bypass, stale-token replay, persisted privileged property, sensitive overexposure, weaker old-version control, or parser-dependent security bypass. Cross-reference a proven workflow issue with access control and dangerous-input findings. Ask whether a single-object issue scales to bulk impact, a skipped approval unlocks a privileged action, or an upload/callback field reaches another trust boundary. Test one safe higher-impact hop only when authorized; do not inflate hypothetical chains. ## Report outcomes and gaps For each finding, preserve the intended invariant, actor/object context, baseline, manipulated request/action, before/after proof, concrete impact, repeatability, cleanup, and faithful reproduction steps. Redact tokens, cookies, payment/customer records, PII, and secrets while retaining field names, value classes, lengths, hashes, and fingerprints. Separate confirmed findings, candidates, expected behavior, unsafe variants not attempted, missing disposable fixtures, untested race conditions, unavailable roles, asynchronous jobs not observed, and other coverage gaps. Never describe incomplete workflow coverage as secure.
Referenced files: 3
tahr-threat-model-app12.4 KB
--- name: tahr-threat-model-app description: Build a full, implementation-backed threat model of an entire existing application, covering actors, assets, trust boundaries, entrypoints, hop-level data flows, abuse cases, connected attack paths, security invariants, control gaps, risk responses, and executable validation handoffs. Use for comprehensive system threat modeling, security architecture assessment, pentest preparation, or correlating a complete application repository with configuration, IaC, API schemas, diagrams, and deployment documentation. Do not use for a feature-only, diff-only, or design-only review. --- # Tahr Threat Model App Model the entire existing application as an attacker would. Use source and configuration for observed implementation, documents for intended behavior, and runtime evidence only when the exact target and test are authorized. Produce a decision and validation plan, not a code-smell list and not a verified-vulnerability report. ## Load the operating contract Before modeling: 1. Read [full-review-workflow.md](references/full-review-workflow.md) for the end-to-end sequence and completion gates. 2. Read [threat-model-ledgers.md](references/threat-model-ledgers.md) for the canonical record relationships and exact enums. 3. Read [threat-evidence-and-quality-gates.md](references/threat-evidence-and-quality-gates.md) before accepting threats, risk ratings, or a complete status. 4. Read [specialist-handoffs.md](references/specialist-handoffs.md) before assigning validation work to another Tahr skill. 5. Use [worked-example.md](references/worked-example.md) only when the expected evidence-to-test trace is unclear. Use [threat-model.template.json](assets/threat-model.template.json) as the starting structure and [threat-model.schema.json](assets/threat-model.schema.json) as the output contract. Do not invent a different report structure. ## Establish full-application scope Record the repository path and revision, packages, services, clients, deployment environments, supplied specifications and documents, excluded third-party internals, runtime authorization, previous model, and unanswered questions. Keep `mode` equal to `full`. Cover every admitted first-party application component. If time, access, or missing evidence prevents full coverage, preserve the complete inventory, disposition each gap, and set `model_status` to `incomplete_high_risk_coverage`. Never silently narrow a full review. Keep these concepts separate: - `model_status`: whether the declared full scope has been dispositioned; - `assurance_status`: whether conclusions are source-observed or also runtime-validated; - `risk`: plausible impact and likelihood of a modeled threat; - `confidence`: strength and completeness of supporting evidence; - `execution_status`: whether a validation test is only planned or has an authorized result. A source-observed model may be `complete` while validation tests remain `planned`, provided runtime validation was not part of the declared scope and all implementation evidence was dispositioned. Express the limitation through `assurance_status`; do not misuse coverage status to imply a test ran. Default to read-only analysis. Do not test production, third parties, real tenants, or real accounts without explicit authorization. Redact credentials, tokens, keys, personal data, and customer data from every artifact. ## Inventory before judging Build a coverage baseline from all first-party files and supplied evidence. Inspect source, manifests and lockfiles, configuration, CI/CD, IaC, containers, API and GraphQL schemas, database models, tests, diagrams, role matrices, workflows, integration notes, and deployment documents. Map: - human, service, worker, administrator, support, peer, tenant, and third-party actors; - critical business, identity, authorization, financial, operational, privacy, audit, and secret assets; - clients, APIs, services, workers, queues, data stores, caches, renderers, control planes, AI systems, and external integrations; - public, internal, administrative, legacy, debug, webhook, job, CLI, upload/download, import/export, socket, serverless, and mobile entrypoints; - authentication, session, authorization, owner/tenant policy, validation, serialization, secrets, logging, rate, quota, and recovery controls; - deployment, network, process, tenant, role, provider, browser, device, and asynchronous trust boundaries. Group routes and files into capabilities and business workflows. Retain exact locations as evidence, but do not turn the report into a route-by-route review. Account for each high-signal item in `coverage`. Freeze a deterministic repository manifest before review. Bind its embedded content hash to the immutable revision and reconcile its admitted paths, packages, environments, documents, and exclusions with metadata. Populate `coverage.inventory.expected_subject_ids` from that inventory before changing any coverage item from `pending`; never shrink the expected set to make a review pass. Create the manifest in the selected threat-model output directory: ```bash python3 <skill-directory>/scripts/build_repository_manifest.py \ path/to/application --include . \ --output path/to/output/repository-manifest.json ``` Add repeated `--include` and `--exclude` arguments when scope is more precise. Pass `--revision` for an immutable VCS/release revision; when omitted, the script derives `snapshot-sha256:<digest>` from the admitted bytes. Copy its embedded revision and content hash (also printed by the script) into metadata, `coverage.inventory.manifest`, and manifest evidence. Use explicit empty arrays for documents or exclusions when there are none; do not omit those scope fields. When the reachable surface is non-trivial and `$tahr-map-attack-surface` is available, send it the frozen revision and scope before deriving threats, then reconcile its inventory into this model rather than treating its output as a second source of truth. If that companion skill is unavailable, perform the same stable inventory locally using this section; do not block or narrow the review. ## Build an evidence-backed graph Record material facts as claim-level evidence. Use exactly `observed`, `intended`, `inferred`, or `unknown`; never combine classes in one field. For every sensitive or state-changing flow, model each hop from the initiating actor to the final response, durable state, or side effect. At each hop record: - source, destination, protocol, input, and affected assets; - actor, user, service, owner, tenant, role, and policy context; - trust boundary crossed; - validation, serialization, authentication, authorization, and logging controls; - queues, workers, callbacks, redirects, repositories, providers, tools, browsers, and other continuation points. Do not stop at a controller when another component performs the security decision. Include data creation, replication, retention, deletion, backup, residency, and third-party handling where material. ## Derive material threats For each critical asset and flow: 1. State the security invariant. 2. Define a realistic actor, goal, preconditions, boundary, abuse steps, affected assets, and business impact. 3. Locate observed and intended controls and their enforcement points. 4. Search for the strongest contradiction in middleware, policy, service, repository, serializer, validator, framework, IaC, or deployment controls. 5. Reject, narrow, or mark the threat `validation_required` according to the contradiction result. 6. Rank risk with an explicit rationale and keep confidence separate. 7. Assign a response, decision, owner, next action, residual risk, and validation test. Use STRIDE as a completeness prompt, not as evidence or a requirement to emit one threat per category. Use ASVS, OWASP API, GraphQL, privacy, mobile, or AI taxonomies only to find omissions and map controls. Retain only threats that connect to the implementation-backed graph. Build attack paths only from connected model IDs. Mark uncertain steps conditional; do not invent a hop merely to make a chain more severe. Keep the canonical model proportional to the application. Reuse a control, invariant, decision, or validation test across related threats when the enforcement point, owner, and discriminating oracle are genuinely the same. Keep claim statements short and reference stable IDs instead of copying the same narrative. Do not create reciprocal records solely to make the artifact look complete. If no candidate survives the evidence, contradiction, and materiality gates, leave `threats`, `attack_paths`, `decisions`, `validation_tests`, and `questions` empty. Preserve the populated inventory, graph, implemented invariants and controls, evidence, coverage, and independent challenge. Never invent a low-value threat or test to avoid an empty ledger. ## Add applicable specialist analysis When personal or regulated data is present, trace collection, linkability, identifiability, detectability, disclosure, consent/awareness, retention, deletion, residency, and third-party processing. When LLMs, agents, RAG, embeddings, MCP, prompts, memory, providers, or tools are present, trace prompt injection, retrieval and memory isolation, tool authorization, confused-deputy paths, output trust, provider disclosure, supply-chain changes, evaluation bypass, and wallet/quota abuse. When a lane is not applicable, record `applicable: false` with evidence. Do not invent threats merely to populate a taxonomy. ## Create executable validation handoffs Create a validation test for every material uncertain, missing, or potentially bypassable control. Include the target Tahr skill, authorization required, safe environment, fixtures, preconditions, normal baseline, exact action, expected control, `attacker_case.attacker_success_signal` and `expected_denial_signal`, `control_case.control_success_signal` and `control_failure_signal`, evidence, cleanup, destructive risk, confidence, and current execution status. Treat each test as `planned` until authorized evidence proves otherwise. Never convert a proposed test into a finding. Route the test to the specialist named in [specialist-handoffs.md](references/specialist-handoffs.md). ## Challenge before publishing Run a separate adversarial quality pass after producing the draft and before publishing it. Use an independent subagent when available; otherwise use a fresh, explicitly separate review pass. Give the reviewer the draft model and source evidence, not the desired conclusions. Require the reviewer to challenge missing assets, actors, boundaries, flow hops, contradictory controls, unsupported impact, inflated risk, route-review drift, unhandled documentation, incomplete coverage, weak actions, broken references, and non-executable tests. Record findings and their dispositions in `quality_review.challenge_findings`. Resolve every high-severity review finding or keep the review and model failed. The reviewer must not silently rewrite the model it is judging. ## Validate and render For a durable model, write only to a user-selected output directory or a clearly named `tahr-threat-model-output/` directory. Never modify application source during the review. Validate before presenting the model as final: ```bash python3 <skill-directory>/scripts/validate_threat_model.py \ path/to/threat-model.json --strict ``` Fix validation failures, rerun the independent challenge when material model content changes, and validate again. Then render concise views: ```bash python3 <skill-directory>/scripts/render_threat_model.py \ path/to/threat-model.json --output-dir path/to/output --strict ``` The renderer produces `threat-model.md`, `validation-plan.json`, and `coverage.json` from the canonical model. Do not edit derived files as though they were independent sources of truth. ## Deliver concise-first results Lead with: 1. model and assurance status; 2. the five most important connected attack paths or threats; 3. blocking security decisions and their owners; 4. the first five validation tests; 5. high-risk coverage gaps. Place complete ledgers after that summary or in the canonical JSON artifact. Avoid repeating the same threat narrative in every section. Never conclude that the application is secure. A clean full model means only that no additional material modeled risks were identified within the stated, completed evidence scope. Preserve assumptions, limitations, accepted risks, pending tests, and a review date so future changes can update the model rather than starting over.
Referenced files: 11
tahr-trace-dangerous-inputs5.87 KB
--- name: tahr-trace-dangerous-inputs description: Trace attacker-controlled input through parsing, validation, normalization, storage, and dangerous server or browser sinks, then safely validate exploitability with class-specific proof gates. Use for injection review, source-to-sink analysis, XSS, SQL/NoSQL injection, command or template injection, SSRF, XXE, path traversal, unsafe deserialization, file upload/processing, webhook, CORS/postMessage, or client-side trust-boundary testing. --- # Trace Dangerous Inputs Find complete attacker-controlled data paths. Do not report a dangerous API call, suspicious regex match, reflection, or error without proving reachability and impact. ## Set a safe review mode 1. Identify repository roots, runtime targets, specifications, traffic, authentication context, and authorized scope. 2. Default to static tracing when live testing is not explicitly authorized. 3. Keep runtime probes low volume, non-destructive, and tied to exact discovered operations. Do not test login credential fields, customer objects, real payment/order flows, or unrelated infrastructure. 4. Require a disposable fixture before uploads, persistent content, webhook registration, or other mutations. Define readback and cleanup first. 5. Use harmless unique markers, controlled callbacks, safe owned canaries, and low-impact commands only. Never delete data, establish persistence, dump broad files/databases, scan internal networks, or alter cloud resources. ## Build the source-to-sink map Read [source-sink-matrix.md](references/source-sink-matrix.md). For every candidate path, record: - entrypoint and exact source: path, query, body, nested field, header, cookie, form, multipart metadata, GraphQL variable, message, file, stored value, or browser source; - parsing and canonicalization order, including decoding, type coercion, content-type selection, duplicate parameters, archive/document parsing, and redirects; - validation, allowlisting, authorization, normalization, encoding, parameterization, and sanitization controls; - transformations and trust-boundary hops across services, queues, jobs, databases, caches, templates, browsers, and third parties; - final sink, execution context, and output/render/fetch/readback path. Trace second-order behavior: input stored now may later reach a query, template, browser render, document processor, shell, webhook, or background job. ## Prioritize real sinks Prioritize operations evidenced by code, schemas, JavaScript, or traffic: - query builders and raw SQL/NoSQL/search expressions; - shell/process APIs, dynamic evaluation, deserializers, and template engines; - URL fetchers, redirects, webhooks, integrations, image/document renderers, and import/export jobs; - XML parsers, file/path/archive operations, upload pipelines, and served-content behavior; - HTML/DOM/navigation/eval-like browser sinks, postMessage handlers, and cross-origin data access. Do not send a payload to the application root merely because a field name looks interesting. Preserve the actual method, body, parser, auth/role/tenant context, and provenance. ## Validate progressively For an authorized runtime target: 1. Establish a clean baseline and the operation's real success semantics. 2. Send a unique inert marker to prove the source reaches the expected context. 3. Change one field, encoding, parser, or carrier at a time using a context-specific safe probe. 4. Distinguish filter/WAF behavior from application execution. 5. Follow the result to the final proof surface: database-derived value, command output, callback, safe file canary, browser runtime effect, persisted readback, internal-service response, or served upload behavior. 6. Repeat with a negative control and preserve exact request/action and response/proof artifacts. If runtime proof is unavailable, report the full reachable code path and missing precondition as a candidate or code-level risk. Do not claim confirmed exploitability. ## Apply class-specific proof gates Read [exploitability-proof-gates.md](references/exploitability-proof-gates.md). Enforce the appropriate gate before confirming a finding. In particular: - require database-derived extraction for SQL injection; - require browser/runtime JavaScript execution for XSS; - require exact-sink callback, internal response, metadata, or file/protocol proof for SSRF; - require command output or controlled callback for command/RCE claims; - separate upload acceptance from browser, server, or document-processing exploitation. Keep timing, errors, reflection, status/size differences, accepted files, listener presence, permissive headers, scanner labels, and source-only hypotheses in a candidate ledger until the gate passes. ## Review the control at the right layer For each path, determine whether the defense is structurally correct: - use parameterized APIs instead of blacklist filtering; - validate canonicalized data at the authoritative server boundary; - allowlist URL schemes/hosts and revalidate every redirect and resolved address; - disable unnecessary parser features and unsafe polymorphic deserialization; - isolate file storage and processing, generate server-side names, and serve inertly; - use context-specific output encoding and safe DOM APIs; - enforce origin and message schema checks before acting on cross-window data. Search for variants of the same unsafe pattern across the repository after proving one path. ## Report proof and coverage For every confirmed issue, provide source-to-sink path, exact entrypoint, control failure, safe proof, impact, affected variants, remediation location, and reproduction steps. Redact secrets and PII while retaining field names, types, lengths, fingerprints, hashes, and proof signals. List untraced sources, unexecuted sinks, parser variants, background jobs, browser-only paths, unavailable callbacks, missing disposable fixtures, and other coverage gaps. Do not convert incomplete tracing into a clean result.
Referenced files: 3
tahr-verify-security-fix5.79 KB
--- name: tahr-verify-security-fix description: Retest a security fix in the exact vulnerable context, decide whether the exploit path is closed, and validate secure remediation and regression coverage without breaking legitimate behavior. Use after a vulnerability patch, remediation commit, PR fix, dependency or configuration change, failed security retest, or when developers need proof that a fix is complete rather than a superficial code change. --- # Tahr Verify Security Fix Verify the failed security control and the attacker outcome, not merely the presence of a patch. Preserve the original finding, proof, context, and scope so a different identity, tenant, route, payload, deployment, or render path cannot produce a false pass. ## Establish retest authority and safety 1. Record the original finding ID, exact claim, required proof, original positive evidence, affected revision, fix revision, and supplied patch. 2. Record the original actor/role, owner/tenant/object, endpoint or workflow, method/content type, payload class, state, configuration, render/trigger context, and final impact. 3. Default to source and test inspection. Exercise only an explicitly authorized local or staging target. Do not infer permission for production, third-party, destructive, credential-changing, billing, or broad data tests. 4. Use disposable accounts and fixtures. Preserve protected identities and extract only the minimum proof sample. 5. Redact passwords, cookies, bearer/session tokens, API keys, private keys, reset codes, personal data, and customer data. Retain only type, location, and hash/fingerprint when necessary. Read [exact-context-retest-verdicts.md](references/exact-context-retest-verdicts.md) and create separate candidate, proof, and coverage records before testing. ## Inspect the remediation Trace the original source-to-sink or missing-control path through current code. Identify: - the exact failed control and where the patch now enforces it; - existing project policy, ownership/tenant scope, validator, sanitizer, encoder, parameter binding, allowlist, network guard, or safe API reused; - adjacent route, resolver, service, repository, serializer, worker, webhook, content type, method, redirect, or render path that may bypass the fix; - public behavior, auth semantics, tenant rules, response shape, data invariants, and framework lifecycle that must remain compatible. Search for a concrete contradiction to the fix claim. A new helper, regex, test, or guard name is not proof that the effective path uses it. ## Retest in the exact context 1. Reproduce the original benign baseline. 2. Replay the original exploit or closest safe equivalent using the original actor, ownership/tenant relation, route, content type, state, and trigger. 3. Confirm the expected denial, encoding, validation, scoped result, blocked network/file action, safe query, or non-execution result. 4. Read back durable state or observe the final sink. Do not stop at the first accepting or rejecting layer. 5. Run a positive control proving the legitimate owner, tenant, role, input, or workflow still succeeds. 6. Run class-relevant alternate encodings, methods, content types, object IDs, roles, tenant relationships, render contexts, redirects, or concurrency variants within the safe budget. 7. Test sibling paths that share the changed control. Use fresh, identity-matched sessions and prove object/tenant ownership for authorization retests. A `200`, body-size change, reflection, upload acceptance, tool success flag, timeout, stale session, or WAF response is not a retest verdict by itself. ## Assign an exact verdict Assign one verdict from the reference contract: - `FIX_VERIFIED`; - `FIX_PARTIAL`; - `NOT_FIXED`; - `REGRESSION_INTRODUCED`; - `INCONCLUSIVE`. Support it with finding-local positive evidence, negative evidence, controls tested, contradiction result, exact-context match, variants, limitations, and final basis. Tooling or environment failure alone requires `INCONCLUSIVE`, not a pass or failure. ## Remediate securely when authorized If the user asks only for verification, report the defect and do not edit. If the user also asks to fix it, read [remediation-and-regression-gates.md](references/remediation-and-regression-gates.md) and: 1. Localize only real repository files, symbols, routes, controls, and tests. 2. Repair the failed control at the smallest correct shared layer. 3. Reuse established project security primitives before creating parallel logic. 4. Avoid blacklist/regex-only fixes when parameterized, encoded, scoped, schema-validated, or allowlisted APIs exist. 5. Preserve public behavior except for the intentional rejection of the vulnerable action. 6. Add a regression test that fails on the vulnerable behavior and a positive control for legitimate behavior. 7. Run focused tests first, then feasible lint, type, build, integration, and broader tests. 8. Review the final diff for adjacent bypasses, superficial controls, generated edits, unrelated churn, secret leakage, and compatibility breaks. Do not mark the fix ready because a test was added; confirm the test reaches the original sink or protected action and fails without the effective control. ## Account for coverage and conclude Account for the original path, every material alternate path, sibling consumer, security-relevant variant, required negative control, legitimate positive control, and test command. Give unresolved high-risk rows an exact blocker and owner. Do not issue `FIX_VERIFIED`, “ready,” or a clean conclusion while an original proof element is missing, the exact context was not reproduced, a high-risk bypass path is unreviewed, required tests did not run, or a material regression remains. Use `FIX_PARTIAL`, `REGRESSION_INTRODUCED`, or `INCONCLUSIVE` and state the next evidence needed.
Referenced files: 3
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 12:00 UTC
- Collection status
- Collected
plugins_6aab2358993881919183acd471020907
Download listing JSON