{"id":21295,"plugin_id":"plugins_6aad979cf8348191812bd2ac0cce6181","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:52.795Z","digest":"04b0be6c4960f7d6eafa5036fb8c4e4eeacde8aa369f34cfa87e8d9603622cef","against":null,"payload":{"description":"Independently evaluate an AI feature, model workflow, RAG system, tool-using agent, prompt, or model/prompt revision with realistic behavioral tasks, grounding, hallucination, unknown/refusal behavior, tool trajectories, adversarial safety cases, latency/cost where material, and regression evidence. Use when AI behavior and quality are the primary object of evaluation. Do not use to design the agent architecture, build broad test infrastructure, choose a provider, or implement repairs.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":277}],"name":"evaluate-ai-systems","skill_md_contents":"---\nname: evaluate-ai-systems\ndescription: Independently evaluate an AI feature, model workflow, RAG system, tool-using agent, prompt, or model/prompt revision with realistic behavioral tasks, grounding, hallucination, unknown/refusal behavior, tool trajectories, adversarial safety cases, latency/cost where material, and regression evidence. Use when AI behavior and quality are the primary object of evaluation. Do not use to design the agent architecture, build broad test infrastructure, choose a provider, or implement repairs.\n---\n\n# Evaluate AI Systems\n\nOwn behavioral AI evaluation, not agent architecture or general test-platform engineering. Evaluate the real candidate against an approved task and safety contract; do not improve prompts or code inside the independent gate.\n\n## Evaluation contract\n\nDefine before execution:\n\n- `TASK SET`: representative normal, edge, ambiguous, unknown, conflicting and high-impact tasks with provenance and version;\n- `EXPECTED OUTCOME`: observable success criteria, allowed variation, required evidence and prohibited outcomes;\n- `EVALUATOR`: deterministic oracle, rubric, qualified human, independent model judge or combination, including limitations and conflict handling;\n- `FAILURE CLASS`: taxonomy such as grounding, retrieval, factuality, hallucination, instruction following, refusal, fallback, tool selection, argument, order, side effect, privacy or safety;\n- `SAFETY SET`: use-case-specific harmful, adversarial, injection, exfiltration, permission and consequential-action cases;\n- `REGRESSION SET`: protected cases, prompt/model/tool/retrieval versions and prior accepted baseline;\n- `PASS / FAIL`: per-criterion threshold and overall gate logic;\n- `STOP CRITERIA`: failure, uncertainty, unsafe behavior, budget exhaustion, insufficient evidence or required owner/specialist escalation.\n\nApply the shared [Product Quality Principles](../../shared/expert-system/product-quality-principles.md), [risk-adaptive assurance](../../shared/expert-system/risk-adaptive-assurance-model.md) and [decision authority](../../shared/expert-system/decision-authority-model.md). The evaluator must not silently change the candidate, expected truth or acceptance threshold.\n\n## Execute proportionately\n\nTest as applicable:\n\n- prompt/model behavior across realistic and held-out scenarios;\n- grounding, source attribution, retrieval relevance/coverage and behavior when evidence is absent or conflicting;\n- factuality, hallucination, calibrated `UNKNOWN`, refusal and safe fallback;\n- tool choice, arguments, ordering, retries, duplicate calls, permission boundaries and consequential actions using safe stubs or isolated environments where possible;\n- indirect/direct prompt injection, untrusted context, data exfiltration, cross-user leakage and instruction-priority conflicts;\n- human control, confirmation, cancellation, recovery and observable failure;\n- response quality, consistency and bias/harm criteria appropriate to audience and use case;\n- latency, cost and resource use only where approved thresholds make them material;\n- prompt, model, retrieval, tool and policy version regression.\n\nKeep evaluation data separate from candidate-authored expected truth. Protect secrets and personal data. Do not execute real purchases, publications, messages, deletions or other consequential side effects merely to test a trajectory.\n\n## Routing boundary\n\n- Agent objectives, tools, permissions, memory, state and architecture → `$design-ai-agents`.\n- Durable CI, fixtures, runners and broad cross-system regression infrastructure → `$engineer-test-and-regression-systems`.\n- Provider/tool selection → `$design-optimal-ai-workflow`.\n- Defect diagnosis or repair → `$diagnose-and-fix-software-defects` or the responsible implementation skill.\n\n## Output\n\nReturn candidate identity and environment, versioned task/safety/regression sets, evaluator design and limitations, raw-result pointers, aggregate and per-class results, failures with reproducible evidence, latency/cost evidence where applicable, regression comparison, verdict `PASS`, `PASS WITH RISKS`, `FAIL`, or `NOT VERIFIABLE`, stop/escalation conditions and exact re-evaluation scope. Never claim universal safety, absence of hallucination or production readiness from a bounded evaluation.\n\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}