← Plugin EvalCONTENT HISTORY

Update to Plugin Eval

Snapshot Sep 30, 2026 · 23:18 UTC · version 0.1.2

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "evaluate-skill",
  "description": "Evaluate a local Codex skill in engineer-friendly terms. Use when the user says \"evaluate this skill\", \"give me an analysis of the game dev skill\", \"audit this skill\", \"why did this score that way\", \"what should I fix first\", or asks for a skill-specific report before benchmarking it.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 96
    }
  ],
  "skill_md_contents": "---\nname: evaluate-skill\ndescription: Evaluate a local Codex skill in engineer-friendly terms. Use when the user says \"evaluate this skill\", \"give me an analysis of the game dev skill\", \"audit this skill\", \"why did this score that way\", \"what should I fix first\", or asks for a skill-specific report before benchmarking it.\n---\n\n# Evaluate Skill\n\nUse this skill when the target is a local skill directory or `SKILL.md` file.\n\n## Workflow\n\n1. Treat \"Evaluate this skill.\" as the default entrypoint.\n2. If the user names a skill instead of giving a path, resolve it locally first, preferring `~/.codex/skills/<skill-name>` and then repo-local `skills/<skill-name>`.\n3. If the user says the request in natural language first, use `plugin-eval start <skill-path> --request \"<user request>\" --format markdown` to show the routed path clearly.\n4. Run `plugin-eval analyze <skill-path> --format markdown`.\n5. Review `At a Glance`, `Why It Matters`, `Fix First`, and `Recommended Next Step` before drilling into details.\n6. Explain which findings are structural, which are budget-related, and which are code-related.\n7. If the user asks for an \"analysis\" of the skill, do not stop at the report. Also run `plugin-eval init-benchmark <skill-path>` and show the setup questions for refining the starter scenarios in `.plugin-eval/benchmark.json`.\n8. If the user wants real usage numbers, switch to \"Measure the real token usage of this skill.\" and run the benchmark flow.\n9. After observed usage is available, use `plugin-eval measurement-plan <skill-path> --observed-usage <usage.jsonl> --format markdown` to recommend what to instrument or improve next.\n10. If the user wants a rewrite plan, route to `../improve-skill/SKILL.md`.\n\n## Skill-Specific Priorities\n\n- frontmatter validity\n- `name` and `description` quality\n- progressive disclosure and reference usage\n- broken relative links\n- oversized `SKILL.md` or descriptions\n- helper script quality for TypeScript and Python files\n\n## Chat Requests To Recognize\n\n- `Evaluate this skill.`\n- `Give me an analysis of the game dev skill.`\n- `Audit this skill.`\n- `Why did this skill score that way?`\n- `What should I fix first?`\n- `Measure the real token usage of this skill.`\n\n## Commands\n\n```bash\nplugin-eval start <skill-path> --request \"Evaluate this skill.\" --format markdown\nplugin-eval analyze <skill-path> --format markdown\nplugin-eval explain-budget <skill-path> --format markdown\nplugin-eval measurement-plan <skill-path> --format markdown\nplugin-eval init-benchmark <skill-path>\nplugin-eval benchmark <skill-path> --dry-run\n```\n\n## Reference\n\n- `../../references/chat-first-workflows.md`\n"
}

SHA-256: 2bd863f3c072c0072001de3f699079afe52334672fda28c0ff19fde54aa66691