{"id":24096,"plugin_id":"plugins_6ab39a1738488191bcb6ba5e381be068","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:17:56.681Z","digest":"7db5a13706b46be3fbdf7cac9bd00ad432cc3084ec5666d14018af00b1383ea0","against":null,"payload":{"name":"prompt-evaluation","description":"Build practical tests and rubrics for comparing prompt behavior.","included_files":[],"skill_md_contents":"---\nname: prompt-evaluation\ndescription: Build practical tests and rubrics for comparing prompt behavior.\n---\n\n# Prompt Evaluation\n\nApply when the user asks whether a prompt works, how to compare versions, or how to test regressions.\n\n1. Translate desired behavior into measurable criteria and identify failure modes.\n2. Create a compact, representative set covering ordinary, ambiguous, boundary, adversarial, and missing-input cases relevant to the task.\n3. Define a scoring rubric with observable anchors; separate factual correctness, instruction following, format, usefulness, and safety.\n4. Recommend a controlled comparison: same model/version, inputs, settings, and tools where possible; repeat stochastic tests and retain outputs.\n5. Treat model-based judging as fallible. Prefer objective checks where feasible and human-review high-impact or subjective outcomes.\n6. Report observed results separately from hypotheses. Record model/date/configuration and rerun tests after material prompt or model changes.\n\nReturn test cases, expected behavior, rubric, and limitations. Never describe a small test set as proof of universal reliability.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}