← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "kermt-monitor",
"description": "Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).",
"included_files": [],
"skill_md_contents": "---\nname: kermt-monitor\ndescription: Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).\nlicense: Apache-2.0\ncompatibility: Requires docker and jq. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n owner: evax@nvidia.com\n classification: atomic-skill\n risk_tier: skill\n# Line/token budget: this file is targeted at ~120 lines / ~1500 tokens — well\n# within the 500-line / 5000-token cap for skill files.\n---\n\n# kermt-monitor\n\nCompanion skill for any KERMT workflow that runs detached: the three pretrain\nskills (`kermt-continue-pretrain`, `kermt-pretrain-scratch`,\n`kermt-add-cmim-pretrain`) plus `kermt-finetune`. `kermt-infer` and\n`kermt-embed` run blocking by default and don't need this skill, but if a\nuser launches them detached on purpose the monitor still works (the\nworkflow-dispatch in step 4 handles unknown workflows by tailing the\nmost-recent log file in the run dir). Reads the run directory's `run.json`,\nqueries docker for the container's state, surfaces the latest progress,\nand either tails or follows the log.\n\n## Hardware requirements\n\nNone. This skill only reads disk + queries docker; no GPU compute.\n\n## Inputs\n\nOne of:\n\n- `<run-dir>` — a positional argument pointing at the directory containing\n `run.json` (e.g. `runs/continue-pretrain_2026-05-17T10-23Z`). Preferred.\n- `--container <name-or-id>` — direct container reference; the skill still\n reads `run.json` from the run dir referenced inside the container's\n inspect output if available, but works degraded-mode without it.\n\nOptional:\n\n- `--lines N` — number of trailing log lines to print (default 50).\n- `--follow` — stream `docker logs -f` until ^C. Useful for \"watch the\n loss\". Without it, the skill is one-shot and exits.\n- `--json` — emit a structured status report instead of human-readable text.\n Useful when the parent agent wants to take downstream action.\n\n## Workflow\n\nLet `RUN_DIR=$1` (or whatever path the user supplies).\n\n1. **Locate the manifest.**\n ```\n MANIFEST=$RUN_DIR/run.json\n ```\n Refuse to proceed if it doesn't exist; surface a helpful message\n pointing the user at the run-dir convention (`runs/<workflow>_<ts>/`).\n\n2. **Parse the manifest** (Python helper):\n ```\n workflow=$(jq -r .workflow $MANIFEST)\n container_name=... # not directly in run.json today; the skill that\n # launched stored it in run.json under\n # container.name during launch (see below note).\n logs_dir=$(jq -r .logs_dir $MANIFEST)\n image_tag=$(jq -r .container.image_tag $MANIFEST)\n started_at=$(jq -r .started_at $MANIFEST)\n ```\n\n3. **Query docker for container state.**\n ```\n docker ps --filter \"name=$container_name\" --format \\\n '{{.ID}}\\t{{.Status}}\\t{{.CreatedAt}}'\n ```\n If absent, fall back to `docker inspect $container_name --format\n '{{.State.Status}} (exit {{.State.ExitCode}})'` to see whether the\n container exited (ok or failed) or was removed (`--rm` after exit).\n\n4. **Find the live log file.**\n ```\n case \"$workflow\" in\n continue-pretrain|pretrain-scratch) LOG=$logs_dir/pretrain_ddp.log ;;\n finetune) LOG=$logs_dir/finetune.log ;;\n *) LOG=$(ls -1t $logs_dir/*.log 2>/dev/null | head -n 1) ;;\n esac\n ```\n The manifest's `workflow` field disambiguates pretrain (`pretrain_ddp.log`)\n from finetune (`finetune.log`). Other workflows fall back to the\n most-recently-modified `.log` in `$logs_dir`.\n\n5. **Show the latest progress.**\n - `tail -n $LINES $LOG` for the raw recent output.\n - Parse the last few progress lines and surface a human-friendly\n summary. The format differs per workflow:\n - Pretrain: epoch / step / val_loss\n ```\n Current epoch: 12/100 step: 4523/9000 val_loss: 0.832 (best 0.821 @ step 4100)\n ```\n - Finetune: fold / epoch / val_<metric> (e.g. val_mae for regression,\n val_auc for classification — read `args_applied.metric` from run.json)\n ```\n Fold 0 epoch 12/30 val_mae 0.187 (best 0.182 @ epoch 9)\n ```\n ```\n Wall-clock: 1h 23m since started_at; ETA ~6h remaining.\n ```\n\n6. **Final test-metrics block (finetune, on completion).** If `workflow` is\n `finetune` AND the container has exited cleanly (`State.Status=exited`,\n `ExitCode=0`) AND `$RUN_DIR/ckpt/fold_*/test_result.csv` exists, parse it\n and emit a per-task metric table:\n ```\n Final test metrics (per task):\n Target MAE\n HLM_clearance 0.187\n RLM_clearance 0.213\n MDR1-MDCK_efflux 0.241\n solubility_pH6.8 0.156\n ```\n The metric column matches `args_applied.metric` (mae for regression, auc\n for classification, etc.). For multi-fold or ensemble runs, average across\n folds/models and note `± std` if std > 0. Skip silently if no\n `test_result.csv` exists (run incomplete or no test split was emitted).\n\n7. **If `--follow`, stream live logs.**\n ```\n docker logs -f $container_name\n ```\n Wraps until ^C.\n\n8. **Stop / cleanup hints** (printed at end of one-shot mode):\n ```\n To stop: docker stop $container_name\n To remove: docker rm $container_name\n To re-run: `$(jq -r .cmd_replay $MANIFEST)`\n ```\n\n## Hard rules\n\n- **Read-only on the user's data.** Never modify `run.json`, never touch the\n container's checkpoint dir. The monitor only inspects.\n- **Don't kill the container without explicit user instruction.** If the\n user asks to stop, run `docker stop`; if they ask to abandon, leave it\n running and just exit.\n- **Don't pull or modify the kermt image.** The monitor only reads.\n- **JSON output mode is non-interactive.** Skip the \"press ^C to exit\"\n prompts and emit a single JSON document so the parent agent can pipe it.\n\n## Note on container_name plumbing\n\nThe run.json schema as currently written does not yet include the launched\ncontainer name — `kermt_run_detached` prints it to stdout but the runner\nscript doesn't capture it into run.json. The monitor falls back to a\nfilesystem-based lookup: list `runs/<workflow>_*/` directories and match by\nmtime; or accept `--container <name>` explicitly. Follow-up: have the\nlaunching skill record container name into run.json before exiting.\n\n## Output (text mode, default)\n\n```\nKERMT continue-pretrain · runs/continue-pretrain_2026-05-17T10-23Z\n Container : kermt-continue-pretrain-… (Up 1 hour, status: running)\n Image : kermt:latest@sha256:…\n Repo : 2fe00f9 (clean)\n Started : 2026-05-17T10:23:14Z (1h 23m ago)\n Workflow : continue-pretrain, pretrain_mode=hybrid, world_size=2\n\n Latest log (last 50 lines from $LOG):\n [Epoch 12/100] step 4523/9000 loss 0.832 lr 1.2e-4\n [val] step 4100 val_loss 0.821 (new best)\n ...\n\n Progress: epoch 12/100, ~12% done. ETA ~6h.\n TensorBoard: tensorboard --logdir $RUN_DIR/logs/tb\n Replay command: $(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n"
}SHA-256: db6362e875c59cf6da7c876b81f3d423b544f622e1f925e0e41239e3b95a7600