{"id":20976,"plugin_id":"plugins_6aab2358993881919183acd471020907","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:41.785Z","digest":"1e2b955b6d26a9d3843de3a10c9ab6a7d88b20207ed334722612bb9b267671d0","against":null,"payload":{"name":"tahr-test-ai-agents","description":"Test security boundaries in applications that use LLM chat, RAG or vector retrieval, memory, file or URL ingestion, model-rendered output, tool/function calling, MCP, or autonomous agents. Use for source-backed AI feature reviews, authorized local or staging runtime tests, prompt-injection assessments, cross-tenant retrieval checks, agent/tool abuse reviews, and AI resource-control testing.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":236},{"relative_path":"references/methodology.md","size_in_bytes":4859},{"relative_path":"references/proof-gates.md","size_in_bytes":3649}],"skill_md_contents":"---\nname: tahr-test-ai-agents\ndescription: Test security boundaries in applications that use LLM chat, RAG or vector retrieval, memory, file or URL ingestion, model-rendered output, tool/function calling, MCP, or autonomous agents. Use for source-backed AI feature reviews, authorized local or staging runtime tests, prompt-injection assessments, cross-tenant retrieval checks, agent/tool abuse reviews, and AI resource-control testing.\n---\n\n# Tahr Test AI Agents\n\nReview the application as an attacker crossing data, instruction, identity, rendering, and tool boundaries. Treat payload catalogs as aids; make the proof discipline the center of the assessment.\n\n## Establish the boundary\n\n1. Confirm the repository, feature, identities, and runtime target placed in scope.\n2. Default to source-only analysis when authorization for active runtime testing is unclear.\n3. Identify protected accounts, production data, third-party integrations, and actions that must remain read-only.\n4. Do not modify application code, configuration, or deployed state unless the user separately requests remediation.\n5. Use disposable tenants, documents, objects, tools, callback collectors, and marker values for active tests.\n\nNever delete data, transfer value, send real messages, publish content, rotate credentials, change privileges, exhaust a customer budget, or exfiltrate real private data. Bound concurrency, output length, request counts, and cost.\n\n## Build the real attack surface\n\nRead [methodology.md](references/methodology.md) before selecting tests.\n\nTrace model and embedding SDK calls to their actual HTTP, WebSocket, queue, or background-job entry points. Include supporting routes for uploads, knowledge bases, conversations, memory, feedback, model settings, tools, and shared views. When source is incomplete, inspect OpenAPI, client bundles, forms, runtime requests, streaming frames, and error shapes.\n\nConfirm an endpoint only when it accepts AI-shaped input or produces generated, streaming, retrieval, embedding, model, or tool behavior. A route name containing `chat`, a generic JSON response, or a successful status code is not confirmation.\n\nFor each confirmed surface, record:\n\n- endpoint, method, controllable field, and authenticated identity;\n- tenant, conversation, memory, and retrieval context;\n- model/provider and guardrail hints, marked `unknown` when unproven;\n- ingestion formats and retrieval sources;\n- renderer and downstream output sinks;\n- available tools, resources, prompts, and autonomous steps;\n- measured rate, token, output, streaming, and quota behavior;\n- evidence source and confidence.\n\n## Test conditionally\n\nEstablish a benign baseline before attack probes. Run only families whose preconditions exist:\n\n- test direct instruction override against an application-specific control;\n- test indirect injection only through content the application really ingests;\n- test retrieval and memory across validated user or tenant boundaries;\n- test output injection in the actual browser or downstream consumer;\n- test tool and MCP agency at real resource and authorization boundaries;\n- test resource controls with conservative measured probes;\n- attempt chains only from previously observed signals.\n\nVary semantic attack concepts before cosmetic encodings. Record every attempt by endpoint, field, concept, payload hash, response class, and success criterion. For stochastic behavior, reproduce the same concept and boundary at least three times in five attempts unless the original security contract requires a stronger threshold.\n\nWhen blocked, ask at most a few fresh-agent passes for concise new concept axes or plausible chains. Treat that advice as leads only. Never let an adviser, model claim, or payload classification verify a finding.\n\n## Apply proof gates\n\nRead [proof-gates.md](references/proof-gates.md) before promoting or scoring a result.\n\nMaintain three separate collections:\n\n1. **Verified findings** — the required boundary and impact proof exists.\n2. **Candidates** — a concrete signal exists, but state the missing proof and safest next check.\n3. **Coverage** — list tested, partially tested, skipped, inaccessible, degraded, and unsafe-to-test surfaces with reasons.\n\nDo not promote model claims, marker-only obedience, generic policy text, stored-but-unretrieved content, API reflection, tool names, version intelligence, one lucky response, or theoretical cost.\n\n## Handle evidence safely\n\nUse raw secrets or private data only transiently when authorized and necessary. Persist the value class, source, tenant or role context, redacted excerpt, length, fingerprint, endpoint, payload class, and reproduction count. Do not place raw tokens, cookies, system prompts, private documents, credentials, PII, or tool output into findings, screenshots, logs, or reports.\n\nName findings after the demonstrated result, not the attempted technique. Calibrate severity to the proven data, action, resource, tenant, or browser boundary. End with prioritized remediation tied to the actual trust boundary and a concise residual-risk statement.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}