← Files MOOS-IvP SkillsARCHIVED FILE

skills/moos-ivp-harness-builder/references/harness-style.md

6.57 KB · Oct 2, 2026 · 00:34 UTC

↓ Download file

# Harness Style

A test harness runs one or more self-evaluating stem
missions. The stem mission decides `grade=pass` or `grade=fail` through
`pMissionEval`. The harness decides which cases to run, how to patch each temp
copy, and how to publish each case row. For ordinary harnesses, the published
row should be directly presentable:

```text
case=<case_name> grade=<pass|fail> eval=<true|false> <evidence fields...>
```

The harness should normally prepend `case=<case_name>` to the stem mission's
result row. It should synthesize a row only when the runner itself failed before
the mission could report a usable result, for example:

```text
case=<case_name> grade=fail reason=launch_error launch_rc=<rc>
case=<case_name> grade=fail reason=missing_result
```

Do not add a harness-owned `reason=` to ordinary mission failures. If
`pMissionEval` reports `grade=fail`, preserve the mission evidence fields that
explain what failed. Use `reason=` for runner failures unless the mission itself
chooses to emit a compact mission-owned reason.

`case=` is the harness row key. Harness case setup should be explicit in the
case matrix, patch files, fixture files, or stem launch arguments. Preserve
useful mission-owned provenance columns such as `form=`, `mhash=`, and domain
evidence fields, but keep harness variation visible in the case definition.

The stem mission must decide the grade through `pMissionEval`. A target file
field such as `grade_hint`, or a shell wrapper that echoes `grade=`, is not a
self-evaluating mission.

## Directory Layout

Use this layout guidance when creating a new harness directory and wiring it to
the stem mission it runs. Prefer root-level `missions/` and `harnesses/`
directories so the harness is visibly separate from the mission it copies,
patches, and runs. For larger projects with many related missions or harnesses,
group both sides by family:

```text
repo-root/
  scripts/
    moos_scoped_teardown.sh
  missions/
    <family>_missions/
      <stem_mission>/
        README.md
        clean.sh
        launch.sh
        launch_shoreside.sh
        launch_vehicle.sh
        xlaunch.sh
        zlaunch.sh
        meta_shoreside.moos
        meta_vehicle.moos
        meta_vehicle.bhv
        results.txt                  # run output
  harnesses/
    <family>_harnesses/
      HNN-<harness_name>/
        README.md
        zlaunch.sh
        results.txt                  # run output
        NSPATCH.md
        <case-name>-shoreside.xmoos
        <case-name>-vehicle.xmoos
        <case-name>-vehicle.xbhv
```

For small projects, the family grouping level may be omitted:

```text
repo-root/
  scripts/
    moos_scoped_teardown.sh
  missions/
    <stem_mission>/
      README.md
      clean.sh
      launch.sh
      launch_shoreside.sh
      launch_vehicle.sh
      xlaunch.sh
      zlaunch.sh
      meta_shoreside.moos
      meta_vehicle.moos
      meta_vehicle.bhv
  harnesses/
    <harness_name>/
      README.md
      zlaunch.sh
      results.txt
```

Use explicit relative paths from the harness to the stem mission. For a harness
under `harnesses/<family>_harnesses/HNN-<harness_name>/`, compute the repository
root with `../../..`; for `harnesses/<harness_name>/`, use `../..`.

Treat the tree as a shape reference, not a mandatory file list. Single-community
unit stems that do not launch a vehicle are the edge case that may omit vehicle
launchers and vehicle templates. Most moving or multi-community stems should
include `launch_shoreside.sh`, `launch_vehicle.sh`, `meta_vehicle.moos`, and
`meta_vehicle.bhv`. Patch-driven harnesses commonly carry `.xmoos`, `.xbhv`, or
`NSPATCH.md`; pure runner harnesses may only need `README.md` and `zlaunch.sh`.
Mission-specific extras such as `plugs.moos`, `case_*.moos`, field files,
obstacle files, or `init_field.sh` are fine when the stem mission needs them,
but do not treat them as required harness-layout defaults.

## Harness Owns

- case matrix documentation
- case-to-setup mapping
- per-case temp mission copies
- serial or rolling execution
- per-case port blocks
- result publication
- runner-failure rows
- scoped teardown
- preserved workdirs for debugging

## Stem Mission Owns

- launchable mission files
- explicit startup initialization
- `pMissionEval` verdict
- `results.txt`
- `zlaunch.sh` compatibility with `xlaunch.sh`
- accepting forwarded ports
- forwarding `--shore_mport`, `--veh_mport`, `--shore_pshare`, and
  `--veh_pshare` from top-level launchers into sublaunchers
- using `nsplug -x` so harness-created `.moosx` and `.bhvx` sidecars are
  consumed during target generation
- expected-negative semantics: if a case is supposed to expose a bad condition,
  `pMissionEval` should report `grade=pass` when that bad condition is observed

Keep this separation strict. If harness code starts interpreting raw MOOS
traffic that `pMissionEval` could grade, the boundary is probably wrong. If
stem wrapper code manufactures `results.txt`, fix the stem before building the
harness.

Use `case_result=success|mismatch|error`, `expected=fail actual=fail`, or
similar wrapper verdicts only for legacy compatibility or when the harness is
explicitly testing failure machinery such as `pMissionEval`, `uMayFinish`, or
CLI return-code behavior.

## App-Level Vs Integration Harnesses

App-level harnesses should grade the app under test. Do not make an app-level
case pass or fail on an unrelated mission event such as vehicle arrival unless
arrival is part of the behavior being tested.

Moving/integration harnesses may grade mission outcome, encounter outcome,
arrival, obstacle interaction, or collision state because those outcomes are the
subject of the test.

For H02-style threshold families, prefer geometry-driven below/edge/above cases
with stock behavior parameters held fixed. Avoid defining the threshold by
changing behavior parameters such as `min_util_cpa`, `max_util_cpa`, or a broad
behavior patch. If the split only works by changing the behavior under test, it
is not a stable H02 threshold pattern.

If a harness must inspect a structured payload, keep that check narrow and
supplemental. Prefer a stem mission that normalizes the payload into a boolean
or scalar and writes the mission-owned grade.

For obstacle, contact, or geometry-heavy harnesses, also read the eval mission
builder's `scenario-and-grading.md`. Keep the scenario model realistic enough to
test the app or behavior path being claimed, and avoid copying reference
geometry blindly when a simpler or safer baseline exposes the same contract.

If a local reference mission's README prose disagrees with its launch scripts,
generated target files, or live results, trust the executable mission evidence
over stale prose.

SHA-256: d2250d928b8493b60b029522feba3aa5967098b89574a27e580ebea2bbf94ead