← Plugin EvalCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Plugin Eval
Snapshot Sep 30, 2026 · 23:18 UTC · version 0.1.2
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "evaluate-plugin",
"description": "Evaluate a local Codex plugin in engineer-friendly language. Use when the user says \"evaluate this plugin\", \"audit this plugin\", \"why did this score that way\", \"what should I fix first\", \"help me benchmark this plugin\", or asks for a plugin-wide report before comparing versions.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 98
}
],
"skill_md_contents": "---\nname: evaluate-plugin\ndescription: Evaluate a local Codex plugin in engineer-friendly language. Use when the user says \"evaluate this plugin\", \"audit this plugin\", \"why did this score that way\", \"what should I fix first\", \"help me benchmark this plugin\", or asks for a plugin-wide report before comparing versions.\n---\n\n# Evaluate Plugin\n\nUse this skill when the target is a plugin root with `.codex-plugin/plugin.json`.\n\n## Workflow\n\n1. Treat \"Evaluate this plugin.\" as the default entrypoint.\n2. If the request comes in as natural chat language, use `plugin-eval start <plugin-root> --request \"<user request>\" --format markdown` first so the user sees the routed local path.\n3. Run `plugin-eval analyze <plugin-root> --format markdown`.\n4. Read `Fix First` before drilling into manifest findings, nested skill findings, and code or coverage details.\n5. If the plugin contains multiple skills, summarize the strongest and weakest ones explicitly.\n6. If the user wants measured usage, switch to \"Help me benchmark this plugin.\" and use the starter benchmark flow.\n7. If the user wants trend data, compare two JSON outputs with `plugin-eval compare`.\n\n## Chat Requests To Recognize\n\n- `Evaluate this plugin.`\n- `Audit this plugin.`\n- `Why did this score that way?`\n- `What should I fix first?`\n- `Help me benchmark this plugin.`\n- `What should I run next?`\n\n## Commands\n\n```bash\nplugin-eval start <plugin-root> --request \"Evaluate this plugin.\" --format markdown\nplugin-eval analyze <plugin-root> --format markdown\nplugin-eval start <plugin-root> --request \"What should I run next?\" --format markdown\nplugin-eval compare before.json after.json\nplugin-eval report result.json --format html --output ./plugin-eval-report.html\nplugin-eval init-benchmark <plugin-root>\nplugin-eval benchmark <plugin-root> --dry-run\n```\n\n## Reference\n\n- `../../references/chat-first-workflows.md`\n"
}SHA-256: 246b35f0a1c0f03b0624982a385e1059116fd0f8fc77bfcf2f550b8cc0e64270