{"id":21292,"plugin_id":"plugins_6aadb63c78b88191a10493870d2f5f6d","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:52.759Z","digest":"1f36c45108ac0a7b07055ed9883c36a8d51defe7c14d91a6f9191647b0d43a19","against":null,"payload":{"description":"Build or repair one self-evaluating MOOS-IvP mission by adding an evaluation layer to an ordinary mission: headless startup, pMissionEval grading, results.txt, uMayFinish/xlaunch completion, zlaunch automation, scoped teardown, and single-scenario validation.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":345},{"relative_path":"assets/eval-single-vehicle/README.md","size_in_bytes":2203},{"relative_path":"assets/eval-single-vehicle/clean.sh","size_in_bytes":1179},{"relative_path":"assets/eval-single-vehicle/launch.sh","size_in_bytes":6986},{"relative_path":"assets/eval-single-vehicle/launch_shoreside.sh","size_in_bytes":4634},{"relative_path":"assets/eval-single-vehicle/launch_vehicle.sh","size_in_bytes":5798},{"relative_path":"assets/eval-single-vehicle/meta_shoreside.moos","size_in_bytes":5507},{"relative_path":"assets/eval-single-vehicle/meta_vehicle.bhv","size_in_bytes":1196},{"relative_path":"assets/eval-single-vehicle/meta_vehicle.moos","size_in_bytes":4183},{"relative_path":"assets/eval-single-vehicle/plug_origin_warp.moos","size_in_bytes":75},{"relative_path":"assets/eval-single-vehicle/zlaunch.sh","size_in_bytes":3598},{"relative_path":"assets/moos-ivp-logo.png","size_in_bytes":1328624},{"relative_path":"assets/moos_scoped_teardown.sh","size_in_bytes":8247},{"relative_path":"references/eval-mission-style.md","size_in_bytes":3352},{"relative_path":"references/evaluator-apps.md","size_in_bytes":6305},{"relative_path":"references/scenario-and-grading.md","size_in_bytes":2106},{"relative_path":"references/validation.md","size_in_bytes":3103},{"relative_path":"references/zlaunch-xlaunch.md","size_in_bytes":3604},{"relative_path":"scripts/live_check_eval_mission.sh","size_in_bytes":6819},{"relative_path":"scripts/static_check_eval_mission.sh","size_in_bytes":5155}],"name":"moos-ivp-eval-mission-builder","skill_md_contents":"---\nname: moos-ivp-eval-mission-builder\ndescription: \"Build or repair one self-evaluating MOOS-IvP mission by adding an evaluation layer to an ordinary mission: headless startup, pMissionEval grading, results.txt, uMayFinish/xlaunch completion, zlaunch automation, scoped teardown, and single-scenario validation.\"\n---\n\n# MOOS-IvP Eval Mission Builder\n\n## Overview\n\nUse this skill for one self-evaluating mission folder: a normal MOOS-IvP mission\nwith an added single-run grading contract. The mission should still be readable\nand runnable by a person, but it must also run headlessly, decide pass/fail\ninside the mission, write `results.txt`, and finish through the shared\n`xlaunch.sh` / `uMayFinish` path.\n\nFor ordinary mission layout, use `moos-ivp-mission-builder` first. For multi-case\nmatrices, patch sweeps, parallel runs, or expected-vs-actual aggregation, use\n`moos-ivp-harness-builder`. For post-run `.alog` evidence, use\n`moos-alog-analysis`.\n\n## Core Rules\n\n- Start from an ordinary mission that already launches cleanly. Prefer the\n  `moos-ivp-mission-builder` baselines or an existing nearby mission family.\n- Add only the evaluation plumbing needed for one scenario:\n  optional `pAutoPoke`, optional `uTimerScript`, `pMissionEval`,\n  `results.txt`, and a thin `zlaunch.sh`.\n- Keep `launch.sh` human-facing. It may accept `--xlaunched`, `--nogui`, and\n  port overrides, but it should not contain case loops or result aggregation.\n- Keep `zlaunch.sh` thin: parse automation arguments, truncate `results.txt`,\n  call shared `xlaunch.sh`, validate that `results.txt` contains `grade=`, then\n  apply project-local scoped cleanup.\n- Let `xlaunch.sh --max_time=<secs>` own `uMayFinish` and the timed wait/stop\n  contract. Do not duplicate that lifecycle in mission-local wrappers.\n- Do not synthesize `grade=` or write the final result row from `zlaunch.sh`,\n  `launch.sh`, or target-file parsing. `pMissionEval` must own the verdict and\n  write `results.txt`; wrappers may only truncate, launch, wait, validate\n  presence of `grade=`, and clean up.\n- For cleanup backstops, copy `assets/moos_scoped_teardown.sh` into the target\n  project as `<project-root>/scripts/moos_scoped_teardown.sh` if it does not\n  already exist. Reuse an existing project-root helper unless it is clearly\n  stale or incompatible.\n- Prefer `pAutoPoke` to seed deploy and evaluation variables in moving\n  missions. Unit-style evals may use `uTimerScript` or the app under test for\n  readiness when there is no vehicle/deploy lifecycle. Do not put pass/fail\n  logic in `pAutoPoke`.\n- Use `pMissionEval` as the primary verdict owner. Prefer mission-level booleans\n  or simple scalar checks over harness-side parsing of raw MOOS traffic.\n- Prefer event-driven `pMissionEval` leads: evaluate when the mission-owned\n  completion event occurs. Use `uMayFinish` through `xlaunch.sh --max_time` as\n  the outer infrastructure ceiling. Use a time-driven evaluation-window lead\n  only when non-completion is an expected mission outcome that should produce\n  mission-owned `grade=fail`.\n- Multiple `lead_condition` lines in the same aspect are allowed, but they are\n  ANDed: all must be true before pass/fail conditions are evaluated. A\n  `lead_condition` after pass/fail conditions starts the next ordered aspect.\n  A single `lead_condition` may use textual `or` when each operand is\n  parenthesized, for example `(EVENT_A = true) or (EVENT_B = true)`. Do not use\n  `||`; it is not a supported `LogicCondition` operator.\n- Treat `BHV_ERROR_SEEN=false` as a normal safety/integrity pass condition.\n  Treat `BHV_WARNING` as advisory development evidence by default: inspect and\n  investigate it with appcasts or `.alog` tools, but do not add a sticky\n  `BHV_WARNING_SEEN` mailflag, result column, or pass condition unless the\n  scenario is explicitly warning-intolerant and the warning signal is known to be\n  stable rather than transient/retracted.\n- Keep `results.txt` scalar and parseable. The only hard schema requirement is\n  `grade=<pass|fail>`; fields such as `form=`, `eval=`, `timeout=`, domain\n  facts, and `mhash=` are recommended evidence, not a mandatory metric set.\n- `mission_mod` is optional mission-owned provenance. Use it only when one\n  mission folder intentionally supports multiple named standalone modes. Omit\n  it from single-scenario eval missions, and do not use it to represent harness\n  cases.\n- If a vehicle-local variable is graded shoreside, bridge it explicitly through\n  the vehicle broker and shoreside broker.\n- For GUI-capable eval missions, keep normal operator buttons available. Do not\n  force appcast/realmcast viewer modes unless the evaluation scenario needs it.\n- Do not add `--case`, `--jobs`, temp mission copies, per-case port blocks, or\n  expected-vs-actual aggregation here. Those belong to the harness builder.\n\n## Workflow\n\n1. Confirm the base mission launches and generates targets.\n2. Identify the smallest mission-owned pass/fail signal.\n   - unit-style app variable\n   - behavior end flag\n   - arrival/collision/encounter outcome\n   - load/process/host info signal\n3. Add evaluation state to the relevant `.bhv` or app config.\n   - When adapting an ordinary waypoint mission, make the graded behavior\n     finite, such as `repeat = 0`, or add an explicit completion flag. A\n     repeating operator survey is usually not a valid eval completion signal.\n4. Bridge graded vehicle-local variables to shoreside when needed.\n5. Add `pAutoPoke` or an equivalent explicit initializer for deploy and\n   evaluation variables.\n6. Add `pMissionEval` with simple lead condition(s), clear pass conditions,\n   `result_flag = MISSION_EVALUATED = true`, and `report_file = results.txt`.\n7. Add or update `zlaunch.sh` to set a mission-appropriate `MAX_TIME` default,\n   accept `--max_time=<secs>` as an override, and forward the final value to\n   `xlaunch.sh --max_time=<secs>`.\n8. Add or update `README.md` with scenario, grading signal, and run commands.\n9. Validate target generation, then run the headless cycle and inspect\n   `results.txt`.\n\n## Reference Use\n\n- Read `references/eval-mission-style.md` for boundaries and file layout.\n- Read `references/evaluator-apps.md` before wiring `pAutoPoke` or\n  `pMissionEval`.\n- Read `references/scenario-and-grading.md` before grading obstacles, contacts,\n  moving/integration outcomes, or structured payloads.\n- Read `references/zlaunch-xlaunch.md` before editing automation wrappers.\n- Read `references/validation.md` before reporting an eval mission as done.\n- Copy `assets/eval-single-vehicle/` when a concrete minimal moving example is\n  useful.\n- Copy `assets/moos_scoped_teardown.sh` into the target project as\n  `<project-root>/scripts/moos_scoped_teardown.sh` when the project does not\n  already have an equivalent root-scoped helper.\n- Run `scripts/static_check_eval_mission.sh <mission-dir>` for a quick\n  structural check.\n- Run `scripts/live_check_eval_mission.sh <mission-dir> --port_base=<free-base>`\n  for bundled-example or high-trust validation when MOOS-IvP runtime tools are\n  available.\n- Treat live-check teardown failure as a test failure, show the teardown error,\n  and preserve the temporary workdir for diagnosis.\n\n## Validation Checklist\n\n- `./launch.sh --just_make --nogui <warp>` succeeds.\n- Generated targets contain `pMissionEval`, explicit initialization\n  (`pAutoPoke`, `uTimerScript`, or an app-owned producer), and any evaluator\n  apps needed for reported columns such as `pMissionHash`.\n- If `pMissionHash` is used for `mhash=` evidence, keep it headless-only by\n  default; GUI targets should not launch both `pMissionHash` and\n  `pMarineViewer` unless the overlapping pMarineViewer hash feature is\n  deliberately disabled.\n- Generated targets include bridged graded variables if the verdict depends on\n  vehicle-local posts.\n- `./zlaunch.sh --just_make <warp>` succeeds when `xlaunch.sh` is on `PATH`.\n- Headless `./zlaunch.sh --max_time=<secs> <warp>` exits cleanly.\n- `results.txt` contains one parseable result line with `grade=`.\n- Runtime warnings are investigated during validation; only stable,\n  scenario-relevant warning metrics are surfaced in `results.txt`.\n- High-trust checks use `scripts/live_check_eval_mission.sh` or equivalent to\n  verify result rows, surface warning evidence, and detect leftover listeners on\n  scoped ports.\n- No mission wrapper uses global `ktm`, `pkill`, or unrelated cleanup.\n- Eval wrappers use `<project-root>/scripts/moos_scoped_teardown.sh` as a scoped\n  backstop after `xlaunch.sh`; they do not use global `ktm`, `pkill`, or broad\n  process discovery.\n- GUI runs retain normal operator controls unless the user requested a\n  headless-only mission.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}