← Plugin EvalCONTENT HISTORY

Update to Plugin Eval

Snapshot Sep 30, 2026 · 23:18 UTC · version 0.1.2

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "metric-pack-designer",
  "description": "Design custom metric packs for plugin-eval so teams can add local evaluation rubrics that emit schema-compatible checks and metrics. Use when the user wants their own evaluation criteria or visualizations.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 116
    }
  ],
  "skill_md_contents": "---\nname: metric-pack-designer\ndescription: Design custom metric packs for plugin-eval so teams can add local evaluation rubrics that emit schema-compatible checks and metrics. Use when the user wants their own evaluation criteria or visualizations.\n---\n\n# Metric Pack Designer\n\nUse this skill when the user wants to extend `plugin-eval` with a local rubric.\n\n## Workflow\n\n1. Clarify the custom rubric categories and target kinds.\n2. Define the smallest useful `checks[]` and `metrics[]` payload.\n3. Create a metric-pack manifest plus a script that prints JSON to stdout.\n4. Run the pack through `plugin-eval analyze <path> --metric-pack <manifest.json>`.\n\n## Design Rules\n\n- Keep IDs stable across runs so comparisons stay meaningful.\n- Emit only `checks[]`, `metrics[]`, and optional `artifacts[]`.\n- Do not try to overwrite the core score or summary.\n- Prefer deterministic local signals over subjective text generation.\n\n## Reference\n\n- `../../references/metric-pack-manifest.md`\n"
}

SHA-256: 0cb4916db9daca36ce655fe37f87bfafa3d517c93bb537e6ef61266529f573ab