← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORY

Update to NVIDIA BioNeMo Agent Toolkit

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "kermt-monitor",
  "description": "Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).",
  "included_files": [],
  "skill_md_contents": "---\nname: kermt-monitor\ndescription: Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).\nlicense: Apache-2.0\ncompatibility: Requires docker and jq. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: atomic-skill\n  risk_tier: skill\n# Line/token budget: this file is targeted at ~120 lines / ~1500 tokens — well\n# within the 500-line / 5000-token cap for skill files.\n---\n\n# kermt-monitor\n\nCompanion skill for any KERMT workflow that runs detached: the three pretrain\nskills (`kermt-continue-pretrain`, `kermt-pretrain-scratch`,\n`kermt-add-cmim-pretrain`) plus `kermt-finetune`. `kermt-infer` and\n`kermt-embed` run blocking by default and don't need this skill, but if a\nuser launches them detached on purpose the monitor still works (the\nworkflow-dispatch in step 4 handles unknown workflows by tailing the\nmost-recent log file in the run dir). Reads the run directory's `run.json`,\nqueries docker for the container's state, surfaces the latest progress,\nand either tails or follows the log.\n\n## Hardware requirements\n\nNone. This skill only reads disk + queries docker; no GPU compute.\n\n## Inputs\n\nOne of:\n\n- `<run-dir>` — a positional argument pointing at the directory containing\n  `run.json` (e.g. `runs/continue-pretrain_2026-05-17T10-23Z`). Preferred.\n- `--container <name-or-id>` — direct container reference; the skill still\n  reads `run.json` from the run dir referenced inside the container's\n  inspect output if available, but works degraded-mode without it.\n\nOptional:\n\n- `--lines N` — number of trailing log lines to print (default 50).\n- `--follow` — stream `docker logs -f` until ^C. Useful for \"watch the\n  loss\". Without it, the skill is one-shot and exits.\n- `--json` — emit a structured status report instead of human-readable text.\n  Useful when the parent agent wants to take downstream action.\n\n## Workflow\n\nLet `RUN_DIR=$1` (or whatever path the user supplies).\n\n1. **Locate the manifest.**\n   ```\n   MANIFEST=$RUN_DIR/run.json\n   ```\n   Refuse to proceed if it doesn't exist; surface a helpful message\n   pointing the user at the run-dir convention (`runs/<workflow>_<ts>/`).\n\n2. **Parse the manifest** (Python helper):\n   ```\n   workflow=$(jq -r .workflow $MANIFEST)\n   container_name=...   # not directly in run.json today; the skill that\n                        # launched stored it in run.json under\n                        # container.name during launch (see below note).\n   logs_dir=$(jq -r .logs_dir $MANIFEST)\n   image_tag=$(jq -r .container.image_tag $MANIFEST)\n   started_at=$(jq -r .started_at $MANIFEST)\n   ```\n\n3. **Query docker for container state.**\n   ```\n   docker ps --filter \"name=$container_name\" --format \\\n       '{{.ID}}\\t{{.Status}}\\t{{.CreatedAt}}'\n   ```\n   If absent, fall back to `docker inspect $container_name --format\n   '{{.State.Status}} (exit {{.State.ExitCode}})'` to see whether the\n   container exited (ok or failed) or was removed (`--rm` after exit).\n\n4. **Find the live log file.**\n   ```\n   case \"$workflow\" in\n     continue-pretrain|pretrain-scratch)  LOG=$logs_dir/pretrain_ddp.log ;;\n     finetune)                            LOG=$logs_dir/finetune.log ;;\n     *)                                   LOG=$(ls -1t $logs_dir/*.log 2>/dev/null | head -n 1) ;;\n   esac\n   ```\n   The manifest's `workflow` field disambiguates pretrain (`pretrain_ddp.log`)\n   from finetune (`finetune.log`). Other workflows fall back to the\n   most-recently-modified `.log` in `$logs_dir`.\n\n5. **Show the latest progress.**\n   - `tail -n $LINES $LOG` for the raw recent output.\n   - Parse the last few progress lines and surface a human-friendly\n     summary. The format differs per workflow:\n     - Pretrain: epoch / step / val_loss\n       ```\n       Current epoch: 12/100  step: 4523/9000  val_loss: 0.832 (best 0.821 @ step 4100)\n       ```\n     - Finetune: fold / epoch / val_<metric> (e.g. val_mae for regression,\n       val_auc for classification — read `args_applied.metric` from run.json)\n       ```\n       Fold 0  epoch 12/30  val_mae 0.187 (best 0.182 @ epoch 9)\n       ```\n     ```\n     Wall-clock: 1h 23m since started_at; ETA ~6h remaining.\n     ```\n\n6. **Final test-metrics block (finetune, on completion).** If `workflow` is\n   `finetune` AND the container has exited cleanly (`State.Status=exited`,\n   `ExitCode=0`) AND `$RUN_DIR/ckpt/fold_*/test_result.csv` exists, parse it\n   and emit a per-task metric table:\n   ```\n   Final test metrics (per task):\n     Target              MAE\n     HLM_clearance       0.187\n     RLM_clearance       0.213\n     MDR1-MDCK_efflux    0.241\n     solubility_pH6.8    0.156\n   ```\n   The metric column matches `args_applied.metric` (mae for regression, auc\n   for classification, etc.). For multi-fold or ensemble runs, average across\n   folds/models and note `± std` if std > 0. Skip silently if no\n   `test_result.csv` exists (run incomplete or no test split was emitted).\n\n7. **If `--follow`, stream live logs.**\n   ```\n   docker logs -f $container_name\n   ```\n   Wraps until ^C.\n\n8. **Stop / cleanup hints** (printed at end of one-shot mode):\n   ```\n   To stop:        docker stop $container_name\n   To remove:     docker rm $container_name\n   To re-run:    `$(jq -r .cmd_replay $MANIFEST)`\n   ```\n\n## Hard rules\n\n- **Read-only on the user's data.** Never modify `run.json`, never touch the\n  container's checkpoint dir. The monitor only inspects.\n- **Don't kill the container without explicit user instruction.** If the\n  user asks to stop, run `docker stop`; if they ask to abandon, leave it\n  running and just exit.\n- **Don't pull or modify the kermt image.** The monitor only reads.\n- **JSON output mode is non-interactive.** Skip the \"press ^C to exit\"\n  prompts and emit a single JSON document so the parent agent can pipe it.\n\n## Note on container_name plumbing\n\nThe run.json schema as currently written does not yet include the launched\ncontainer name — `kermt_run_detached` prints it to stdout but the runner\nscript doesn't capture it into run.json. The monitor falls back to a\nfilesystem-based lookup: list `runs/<workflow>_*/` directories and match by\nmtime; or accept `--container <name>` explicitly. Follow-up: have the\nlaunching skill record container name into run.json before exiting.\n\n## Output (text mode, default)\n\n```\nKERMT continue-pretrain · runs/continue-pretrain_2026-05-17T10-23Z\n  Container : kermt-continue-pretrain-…  (Up 1 hour, status: running)\n  Image     : kermt:latest@sha256:…\n  Repo      : 2fe00f9 (clean)\n  Started   : 2026-05-17T10:23:14Z (1h 23m ago)\n  Workflow  : continue-pretrain, pretrain_mode=hybrid, world_size=2\n\n  Latest log (last 50 lines from $LOG):\n    [Epoch 12/100] step 4523/9000 loss 0.832 lr 1.2e-4\n    [val] step 4100 val_loss 0.821 (new best)\n    ...\n\n  Progress: epoch 12/100, ~12% done. ETA ~6h.\n  TensorBoard: tensorboard --logdir $RUN_DIR/logs/tb\n  Replay command: $(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n"
}

SHA-256: db6362e875c59cf6da7c876b81f3d423b544f622e1f925e0e41239e3b95a7600