← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORY

Update to NVIDIA BioNeMo Agent Toolkit

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "kermt-embed",
  "description": "Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (grover_base / cmim / hybrid / finetuned). Writes one .npy per readout type (atom_from_atom, bond_from_atom, atom_from_bond, bond_from_bond) plus canonical_smiles.npy and validity.npy. Calls task/extract_embeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).",
  "included_files": [],
  "skill_md_contents": "---\nname: kermt-embed\ndescription: Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (grover_base / cmim / hybrid / finetuned). Writes one .npy per readout type (atom_from_atom, bond_from_atom, atom_from_bond, bond_from_bond) plus canonical_smiles.npy and validity.npy. Calls task/extract_embeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: workflow-skill\n  risk_tier: skill\n# Line/token budget: targets ~150 lines / ~1800 tokens — within the\n# 500-line / 5000-token cap for skill files.\n---\n\n# kermt-embed\n\nExtract per-molecule embeddings from any encoder-bearing KERMT checkpoint.\nThe skill is the workflow orchestrator: validate ckpt, validate CSV, clean\nSMILES, launch the runner blocking, return the per-readout `.npy` files.\n\n## Hardware requirements\n\n- **GPUs**: 1 (single-GPU).\n- **VRAM**: ≥ 4 GB for the default `batch_size 64`.\n- **Disk**: depends on output size — roughly a few MB per 1k molecules at\n  `hidden 800` per readout, so ~10–20 MB per 1k molecules across the 4\n  readouts. Plus a small `canonical_smiles.npy` + `validity.npy` per run.\n- **Driver / CUDA**: any host supporting CUDA 12.6.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — SMILES CSV. First column is `smiles`; other columns\n  are ignored (no targets needed).\n\nCheckpoint (optional — defaults to the released model if omitted):\n\n- `--ckpt <path>` — any encoder-bearing checkpoint. Grover_base, cmim,\n  hybrid, and finetuned ckpts are all accepted. The validator only refuses\n  ckpts with no encoder. **If omitted**, the skill offers to download the\n  released pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and embed with\n  it — see \"Resolve & validate the checkpoint\" (workflow step 3).\n- `--pretrained-release` — explicit opt-in to use the released model without\n  the interactive prompt (for non-interactive / agent runs). Mutually\n  exclusive with `--ckpt`.\n- `--model-dir <dir>` — where to save the downloaded bundle (default\n  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is\n  reused, not re-downloaded.\n\nOptional:\n\n- `--batch-size N` — override the configured default (64).\n- `--gpus 0` — single GPU id (default 0).\n- `--from-prepare <dir>` — skip the prepare step and reuse an existing\n  `prepare_data.json` in `<dir>`.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout.\n\n1. **Pre-flight: container + system probe.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh check_system\n   ```\n\n2. **Compute run directory.**\n   ```\n   RUN_DIR=$KERMT_REPO/runs/embed_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n   ```\n\n3. **Resolve & validate the checkpoint.**\n\n   **Resolve — only if `--ckpt` was omitted.** Default to the released\n   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:\n   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:\n     \"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2\n     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)\n     and embed with it? [y/N]\". **Never download without an explicit yes** (or\n     `--pretrained-release`). If both `--ckpt` and `--pretrained-release` are\n     given, abort — they conflict.\n   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor\n     `--model-dir <dir>` if given. An already-complete bundle is reused.\n   - **Download** (foreground; ~282 MB on first fetch):\n     ```\n     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \\\n         \"python agent/scripts/fetch_released_model.py --out /model\"\n     ```\n     Parse the JSON; abort on `ok: false` (surface `errors`). On success set\n     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.\n\n   **Validate** the resolved (or user-provided) ckpt:\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \\\n       \"python agent/scripts/check_checkpoint.py --mode embed --ckpt /ckpt\"\n   ```\n   Parse JSON. Abort on `ok: false`. The validator only refuses encoder-less\n   ckpts (rare).\n\n4. **Validate the data.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n       \"python agent/scripts/check_data.py --mode embed --csv /data/<basename>\"\n   ```\n\n5. **Prepare the data** (clean-only — no features step).\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/prepare_data.py --mode embed \\\\\n            --csv /data/<basename> --out /runs/data\"\n   ```\n   Outputs land at `$RUN_DIR/data/prepare_data.json` with a single `clean_csv`\n   path. `task/extract_embeddings.py` featurizes from SMILES on the fly.\n\n6. **Launch the runner (blocking).**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run \\\\\n       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\\\n       \"python agent/scripts/run_extract_embeddings.py \\\\\n            --ckpt /ckpt \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --out /runs \\\\\n            [--gpus 0 --batch-size N]\"\n   ```\n\n7. **Report to the user.**\n   - Embeddings directory: `$RUN_DIR/out/`\n     - `atom_from_atom.npy`, `bond_from_atom.npy`,\n       `atom_from_bond.npy`, `bond_from_bond.npy` (the 4 standard readouts;\n       each shape `(N_rows, hidden_size)`)\n     - `metadata.pkl` — pickle of a dict containing `canonical_smiles`\n       (RDKit-canonicalized SMILES per row), `valid` (boolean per-row: did\n       RDKit parse it), plus other run metadata.\n   - Manifest: `$RUN_DIR/run.json`\n   - Log: `$RUN_DIR/logs/embed.log`\n\n## Hard rules\n\n- **Never download the released model without consent.** When `--ckpt` is\n  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user \"yes\"\n  or an explicit `--pretrained-release` flag. `--ckpt` and\n  `--pretrained-release` are mutually exclusive.\n- **Never modify the user's ckpt.** The runner reads-only via\n  `task/extract_embeddings.py`'s `--checkpoint <path>` flag.\n- **Arch comes from the ckpt.** No `--hidden-size` flag etc. on this runner;\n  `task/extract_embeddings.py` reads arch from the ckpt's saved_args.\n\n## Common errors\n\n- `prepare_data manifest is missing required output 'clean_csv'` → prepare\n  ran with `--skip-clean` but no source CSV given. Re-run prepare without it.\n- `--gpus '0,1' is single-GPU only` → pass a single id.\n\n## Replayability\n\n```bash\n$(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n\nIf `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin\nthe commit via `repo.commit` and `git checkout` it first.\n"
}

SHA-256: a355aa1a6fda8177a1af2582f2ce381cabcf1327ceba4dcd6567329fd86b481c