← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "kermt-infer",
"description": "Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The skill validates that the input ckpt has task FFN heads (refuses pretrain ckpts with a redirect to kermt-finetune), validates the CSV, prepares the data (clean + rdkit_2d features), then launches main.py predict inside the kermt container (blocking, minutes-scale).",
"included_files": [],
"skill_md_contents": "---\nname: kermt-infer\ndescription: Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The skill validates that the input ckpt has task FFN heads (refuses pretrain ckpts with a redirect to kermt-finetune), validates the CSV, prepares the data (clean + rdkit_2d features), then launches main.py predict inside the kermt container (blocking, minutes-scale).\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n owner: evax@nvidia.com\n classification: workflow-skill\n risk_tier: skill\n# Line/token budget: targets ~170 lines / ~2000 tokens — well within the\n# 500-line / 5000-token cap for skill files.\n---\n\n# kermt-infer\n\nRun predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The\nskill is the workflow orchestrator: validate ckpt, validate CSV, prepare data,\nlaunch the runner blocking, return the predictions CSV.\n\n## Hardware requirements\n\n- **GPUs**: 1 (single-GPU). Multi-GPU inference is not currently supported.\n- **VRAM**: ≥ 4 GB for the default `batch_size 32`.\n- **Disk**: a few hundred MB per run (cleaned CSV + features + predictions).\n- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).\n\n## Inputs\n\nRequired:\n\n- `--ckpt <path>` — finetuned checkpoint (must have task FFN heads). The\n validator refuses pretrain ckpts with a redirect to `kermt-finetune`.\n- `--csv <path>` — SMILES-only CSV. First column is `smiles`; other columns\n are ignored.\n\nOptional:\n\n- `--batch-size N` — override the configured default (32).\n- `--seed N` — random seed for inference (deterministic featurization paths).\n- `--gpus 0` — single GPU id (default 0). Multi-GPU rejected.\n- `--from-prepare <dir>` — skip the prepare step and reuse an existing\n `prepare_data.json` in `<dir>`.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout, and assume\n`kermt-setup` has built `kermt:latest`.\n\n1. **Pre-flight: ensure container + system probe.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh check_system\n ```\n Refuse to proceed on `ok: false`.\n\n2. **Compute run directory.**\n ```\n RUN_DIR=$KERMT_REPO/runs/infer_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n ```\n\n3. **Validate the checkpoint.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \\\n \"python agent/scripts/check_checkpoint.py --mode inference --ckpt /ckpt\"\n ```\n Parse the JSON. Abort on `ok: false`. The validator rejects pretrain ckpts\n (`has_task_ffn: false`) with a redirect to `kermt-finetune`.\n\n4. **Validate the data.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n \"python agent/scripts/check_data.py --mode inference --csv /data/<basename>\"\n ```\n Abort on `ok: false`.\n\n5. **Prepare the data.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n \"python agent/scripts/prepare_data.py --mode inference \\\\\n --csv /data/<basename> --out /runs/data\"\n ```\n Outputs land at `$RUN_DIR/data/prepare_data.json` with `clean_csv` +\n `clean_npz` paths (rdkit_2d_normalized features).\n\n6. **Launch the runner (blocking).**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run \\\\\n --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\\\n \"python agent/scripts/run_inference.py \\\\\n --ckpt /ckpt \\\\\n --prepare-manifest /runs/data/prepare_data.json \\\\\n --out /runs \\\\\n [--gpus 0 --batch-size N --seed N]\"\n ```\n Returns the predictions CSV path on success.\n\n7. **Report to the user.** Output a short summary:\n - Predictions: `$RUN_DIR/out/predictions.csv` (smiles + per-target columns)\n - Manifest: `$RUN_DIR/run.json` (cmd_replay + image digest + applied args)\n - Log: `$RUN_DIR/logs/inference.log`\n - Row count: <N> molecules predicted across <K> targets\n\n## Hard rules\n\n- **Never modify the user's ckpt.** The runner symlinks the ckpt into a\n unique `<out>/ckpt_link/` subdir so `main.py predict --checkpoint_dir`\n picks it up; the source file stays untouched.\n- **Arch comes from the ckpt, never from CLI/defaults.** The runner records\n the validator's arch block in `run.json` but does not pass arch flags into\n `main.py predict` — predict reads them from the loaded ckpt's saved_args.\n- **Single-GPU only.** Multi-GPU inference is not currently supported.\n- **Echo applied defaults.** The `args_applied` field of `run.json` records\n every flag's value + source (user / default-config). Surface a short\n summary of any default-filled flag.\n\n## Common errors\n\n- `inference requires a finetuned ckpt with task FFN heads` → ckpt is a\n pretrain ckpt; use `kermt-finetune` first.\n- `prepare_data manifest reports ok=False` → check the manifest `errors` for\n the failed step (typically clean_smiles or save_features).\n- `could not convert string to float: '<value>'` from save_features or main.py\n predict → input CSV has a non-numeric passthrough column (e.g. a 'split'\n label). The prep step now strips the CSV to SMILES-only at inference; if\n this error still surfaces, the CSV is being read by a runner that bypassed\n prepare_data. Re-run via the skill, not `main.py` directly.\n- `--gpus '0,1' is single-GPU only` → pass a single id.\n\n## Replayability\n\nThe `run.json` `cmd_replay` field is a single-line command that re-runs the\ninference with the same inputs. To replay inside the kermt container:\n\n```bash\n$(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n\nIf `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin\nthe commit via `repo.commit` and `git checkout` it first.\n"
}SHA-256: 0bc0b93349f0edc8aa6880cfdafef22abfcc6054edb58c4abf70146d532c0059