← Prompt EngineerCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Prompt Engineer
Snapshot Sep 30, 2026 · 23:17 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "prompt-evaluation",
"description": "Build practical tests and rubrics for comparing prompt behavior.",
"included_files": [],
"skill_md_contents": "---\nname: prompt-evaluation\ndescription: Build practical tests and rubrics for comparing prompt behavior.\n---\n\n# Prompt Evaluation\n\nApply when the user asks whether a prompt works, how to compare versions, or how to test regressions.\n\n1. Translate desired behavior into measurable criteria and identify failure modes.\n2. Create a compact, representative set covering ordinary, ambiguous, boundary, adversarial, and missing-input cases relevant to the task.\n3. Define a scoring rubric with observable anchors; separate factual correctness, instruction following, format, usefulness, and safety.\n4. Recommend a controlled comparison: same model/version, inputs, settings, and tools where possible; repeat stochastic tests and retain outputs.\n5. Treat model-based judging as fallible. Prefer objective checks where feasible and human-review high-impact or subjective outcomes.\n6. Report observed results separately from hypotheses. Record model/date/configuration and rerun tests after material prompt or model changes.\n\nReturn test cases, expected behavior, rubric, and limitations. Never describe a small test set as proof of universal reliability.\n"
}SHA-256: 7db5a13706b46be3fbdf7cac9bd00ad432cc3084ec5666d14018af00b1383ea0