← MOOS-IvP SkillsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to MOOS-IvP Skills
Snapshot Sep 30, 2026 · 23:16 UTC · version 1.4.12
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Build or repair multi-case MOOS-IvP test harnesses around self-evaluating stem missions: case matrices, per-case mission copies, result aggregation, serial and rolling parallel execution, port isolation, scoped teardown, and nspatch variants. Use moos-ivp-eval-mission-builder for stem missions.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 331
},
{
"relative_path": "assets/moos-ivp-logo.png",
"size_in_bytes": 1328624
},
{
"relative_path": "assets/moos_scoped_teardown.sh",
"size_in_bytes": 8247
},
{
"relative_path": "references/case-matrix.md",
"size_in_bytes": 1991
},
{
"relative_path": "references/example-harness-zlaunch.md",
"size_in_bytes": 12284
},
{
"relative_path": "references/generated-harness-self-tests.md",
"size_in_bytes": 7446
},
{
"relative_path": "references/harness-style.md",
"size_in_bytes": 6724
},
{
"relative_path": "references/nspatch-workflow.md",
"size_in_bytes": 2843
},
{
"relative_path": "references/ports-and-parallelism.md",
"size_in_bytes": 3653
},
{
"relative_path": "references/scoped-teardown.md",
"size_in_bytes": 3464
},
{
"relative_path": "references/timing-and-benchmarking.md",
"size_in_bytes": 1685
},
{
"relative_path": "references/validation.md",
"size_in_bytes": 3779
},
{
"relative_path": "scripts/static_check_harness.sh",
"size_in_bytes": 7180
}
],
"name": "moos-ivp-harness-builder",
"skill_md_contents": "---\nname: moos-ivp-harness-builder\ndescription: \"Build or repair multi-case MOOS-IvP test harnesses around self-evaluating stem missions: case matrices, per-case mission copies, result aggregation, serial and rolling parallel execution, port isolation, scoped teardown, and nspatch variants. Use moos-ivp-eval-mission-builder for stem missions.\"\n---\n\n# MOOS-IvP Harness Builder\n\n## Overview\n\nUse this skill for a harness that runs one or more self-evaluating stem missions\nacross multiple named cases. The stem mission should own the mission grade. The\nharness should own case selection, patching, temp copies, port isolation,\nrolling parallel execution, cleanup, and direct publication of per-case result\nrows.\n\nFor the stem mission itself, use `moos-ivp-eval-mission-builder`. For ordinary\nmission construction before evaluation plumbing, use `moos-ivp-mission-builder`.\nFor post-run `.alog` evidence, use `moos-alog-analysis`.\n\n## Core Rules\n\n- Start from a stem mission that runs headlessly and writes `results.txt` with a\n `grade=` column.\n- Prefer placing harness directories at the repository root, alongside\n `missions/`, for example `harnesses/<harness_name>/` paired with\n `missions/<stem_mission>/`. In larger repositories, use optional family\n grouping for both sides, such as\n `harnesses/<family>_harnesses/HNN-<harness_name>/` paired with\n `missions/<family>_missions/<stem_mission>/`. Have the harness refer to stem\n missions with explicit relative paths. Other layouts are acceptable when\n project conventions or packaging require them.\n- The stem mission must be a real eval mission: `pMissionEval` writes the\n `grade=` row. Do not accept a stem where `zlaunch.sh` or shell code\n synthesizes `grade=` from target files, patch markers, or harness knowledge.\n- Keep case intent documented in the harness README under `Cases` or\n `Current Matrix`.\n- Use exact case tokens in documentation and in `zlaunch.sh`.\n- `case=` is the harness-owned variation identity. Harnesses must not set,\n derive, require, or interpret `mmod`; any `mmod=` field produced by the\n mission is opaque mission-owned provenance.\n- Keep case setup explicit. A shell `case` block mapping case name to patch\n files, fixture files, stem launch arguments, and intent is easier to audit\n than filename inference.\n- When multiple cases reuse one stem but need different setup or evaluation\n criteria, express the differences in the case matrix and case-owned patch\n files, fixture files, or stem launch arguments.\n- Keep `launch.sh` and stem wrappers human-facing. Put loops, temp copies,\n aggregation, and archives in harness code.\n- Prefer mission-owned grades. The harness should normally prepend\n `case=<case_name>` to the mission result row and preserve the mission's\n `grade=pass|fail` as the case verdict.\n- For expected-negative cases, make the stem `pMissionEval` pass when the\n expected negative evidence is observed. Do not encode those cases as\n `expected=fail actual=fail` unless the harness is explicitly testing failure\n machinery such as `pMissionEval`, `uMayFinish`, or CLI return semantics.\n- Harness code should synthesize its own `grade=fail` rows only for runner\n failures, such as `reason=launch_error`, `reason=missing_result`,\n `reason=prepare_error`, `reason=missing_result_file`, or\n `reason=teardown_error`.\n- Do not add a harness-owned `reason=` for ordinary `pMissionEval` failures.\n Preserve the mission evidence columns that explain the failure. A mission may\n report its own compact `reason=`, but the harness should not reinterpret it.\n- Avoid new `case_result=success|mismatch|error` result formats for ordinary\n harnesses. Treat them as legacy compatibility or as a special pattern for\n tests whose subject is the failure machinery itself.\n- Keep evaluation levels strict: app-level harnesses should grade the app under\n test; moving/integration harnesses may grade arrival, encounter outcome,\n collision state, or other mission outcomes.\n- Expose `--case`, `--port_base`, `--keep_workdirs`, `--gui`, `--nogui`, and\n `--max_time` when the harness can support them. New generated harnesses must\n expose `--jobs`, default it to `1`, and run real backgrounded cases when it is\n greater than `1`. Prefer Bash 5.1+ rolling scheduling with\n `wait -p <pidvar> -n`, so the next pending case starts as soon as any active\n case finishes. Batch-barrier waves are a legacy compatibility pattern and do\n not satisfy the new generated harness contract.\n- Modern generated harnesses may require Bash 5.1+ for reliable rolling\n scheduling and PID-to-case bookkeeping. Use `#!/usr/bin/env bash`, add an\n early Bash version guard with a clear macOS/Homebrew message, and optionally\n re-exec a known Homebrew/Linuxbrew Bash before failing.\n- Treat harness `--max_time` as a run-time ceiling override forwarded to each\n stem eval mission's `zlaunch.sh`; do not use it as harness-side grading\n logic.\n- Default generated harnesses to `PORT_BASE=9000`. Use higher fresh bases only\n as explicit run-time overrides for automation or local sessions that may\n collide with ordinary missions in the `9000` range.\n- Use headless mode as the default. Keep `--gui` available for an individual\n case when visual inspection is useful.\n- For parallel execution, give each live case its own temp mission copy and\n port block. Do not patch or run through a shared stem directory while\n multiple cases are active.\n- Create per-case temp mission copies under a harness-owned run root, not a\n generic system temp location. `--keep_workdirs` should preserve one auditable\n run tree beneath the harness directory.\n- Use scoped teardown between cases and at harness exit. Prefer\n copying `assets/moos_scoped_teardown.sh` into the generated project as\n `<project-root>/scripts/moos_scoped_teardown.sh`, sourcing it from harness\n launchers, and calling `moos_scoped_teardown_stop_root` on the harness-owned\n run root or case directory. Do not use global `ktm`, `pkill`, or machine-wide\n cleanup as the normal path.\n## Workflow\n\n1. Confirm the stem mission passes as a single eval mission.\n - run the eval mission static checker against the stem\n - `launch.sh` accepts and forwards `--shore_mport`, `--veh_mport`,\n `--shore_pshare`, and `--veh_pshare`.\n - launchers use `nsplug -x` so `.moosx` and `.bhvx` sidecars are consumed.\n - generated targets prove the forwarded ports and patches actually landed.\n2. Define case tokens, case intent, and the mission-owned evidence each case\n should report.\n3. Document the case matrix in `README.md`.\n4. Decide whether each case needs patch files, fixture files, stem launch\n arguments, or no setup changes.\n5. Build `zlaunch.sh` around:\n - argument parsing\n - case selection and setup mapping\n - optional patch overlay application\n - `run_case`\n - serial and rolling execution\n - result aggregation\n - cleanup traps\n6. Add the teardown helper asset to the generated project if there is not\n already an equivalent root-scoped helper.\n7. Implement port forwarding from harness to stem mission and verify generated\n targets reflect those ports.\n8. Add `--keep_workdirs` for debugging preserved temp copies.\n9. Validate one case, `--jobs=1`, then a small rolling run on a fresh\n `--port_base`.\n\n## Reference Use\n\n- Read `references/harness-style.md` for the overall architecture.\n- Read `references/case-matrix.md` before writing README case docs.\n- Read `references/nspatch-workflow.md` before adding patch overlays.\n- Read `references/ports-and-parallelism.md` before implementing `--jobs` or\n `--port_base`.\n- Read `references/generated-harness-self-tests.md` before reporting a new or\n heavily changed harness as trustworthy.\n- Read `references/validation.md` before reporting a harness as done.\n- Read `references/timing-and-benchmarking.md` before tuning `--jobs`, sleeps,\n `--max_time`, or benchmarking rolling runs.\n- Read `references/scoped-teardown.md` before writing cleanup logic.\n- Read `references/example-harness-zlaunch.md` for a compact runner skeleton.\n- Reuse `assets/moos_scoped_teardown.sh` by copying it into generated harness\n projects as `<project-root>/scripts/moos_scoped_teardown.sh` when they do not\n already provide an equivalent root-scoped helper.\n- Run `scripts/static_check_harness.sh <harness-dir>` for a structural check.\n\n## Validation Checklist\n\n- Stem mission passes alone with `./zlaunch.sh --max_time=<secs>`.\n- Stem mission passes `moos-ivp-eval-mission-builder` static validation; the\n harness static checker alone is not enough.\n- Harness README has a `Cases` or `Current Matrix` section with exact case\n tokens and prose intent.\n- `./zlaunch.sh --case=<case> --max_time=<secs>` works for at least one nominal\n case and one expected-negative case if the suite has both.\n- `./zlaunch.sh --jobs=1 --port_base=<base>` works, and a rolling run with\n `--jobs=2` or higher uses distinct temp directories and distinct port blocks.\n New generated harnesses should start the next pending case whenever an active\n case finishes, not wait for an entire batch barrier.\n- Aggregated results include `case=` and the mission's original result columns,\n especially `grade=` and useful evidence fields such as `eval=`,\n `warning_count=`, `expected=`, `observed=`, or case-specific scalars.\n- `case=` is the harness row key. Harness case setup should be explicit in the\n case matrix, patch files, fixture files, or stem launch arguments. `form=`,\n `mhash=`, and mission-owned evidence columns may be preserved as provenance.\n- Ordinary case success is `grade=pass`. Any row with `grade!=pass` should make\n the harness exit nonzero unless the harness is explicitly testing failure\n machinery.\n- Harness-owned failure rows use `case=<case> grade=fail reason=<runner_reason>`\n and preserve launch return codes or setup evidence when available.\n- Selected runs produce one normalized result line for every selected case,\n including setup errors and intentional failures.\n- A selected run that produces zero case rows is a harness failure and should\n exit nonzero with a clear diagnostic. This catches portability bugs where the\n case loop never actually ran.\n- New generated harnesses that implement rolling scheduling should require Bash\n 5.1+ and check that requirement near the top of `zlaunch.sh`. For legacy\n portable harnesses that intentionally target macOS system Bash 3.2, avoid\n `mapfile`, `readarray`, associative arrays, `wait -n`, and `wait -p`.\n- `--keep_workdirs` preserves enough files to inspect generated targets and\n `results.txt`.\n- Preserved workdirs show generated targets using distinct forwarded ports and\n any intended `.moosx` / `.bhvx` sidecars.\n- No harness path relies on global `ktm`, `pkill`, or `killall`.\n- Harness cleanup uses a root-scoped teardown helper or an equivalent recorded\n PID cleanup path; generated harnesses should not invent broad process cleanup.\n- A teardown failure is visible, makes an otherwise successful run fail, and\n preserves the affected run root for inspection.\n- Logs do not contain unexpected warnings hidden by case aggregation.\n"
}SHA-256 of public snapshot: dcec88a905ddb5c1ed922131f86d56933550e2dc23d8b93ea4e75e4d810044b7