← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "kermt-add-cmim-pretrain",
"description": "Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.",
"included_files": [],
"skill_md_contents": "---\nname: kermt-add-cmim-pretrain\ndescription: Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n owner: evax@nvidia.com\n classification: workflow-skill\n risk_tier: skill\n# Line/token budget: ~165 lines, ~1900 tokens — well within the\n# 500-line / 5000-token cap for skill files.\n---\n\n# kermt-add-cmim-pretrain\n\nConvert a grover_base checkpoint (legacy original-GROVER `grover.encoders.*`\nor modern `kermt.encoders.*`, with or without vocab heads) into a fully-formed\nhybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's\ncorpus as hybrid.\n\nThis is a thin wrapper: `upgrade_to_hybrid.py` produces a new ckpt that\nclassifies as `model_type: hybrid` via `check_checkpoint.py`, and the rest of\nthe workflow is identical to `kermt-continue-pretrain`.\n\n> **Status: experimental.** This workflow is functional end-to-end but has not\n> been benchmarked against the manuscript's from-scratch hybrid training (which\n> produces the released checkpoint). Use as an experimental alternative to\n> `kermt-pretrain-scratch` when you want to extend an existing grover_base\n> checkpoint rather than restart from random init. Validate downstream\n> performance on your own benchmark before relying on the upgraded ckpt for\n> production work.\n\n## Hardware requirements\n\nSame as `kermt-continue-pretrain` (the cMIM decoder adds parameters but not\nsubstantially; VRAM headroom should be fine). The upgrade step itself is\nfast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes\nGPU.\n\n## When to invoke\n\n- User has a grover_base checkpoint (encoder-only or with vocab heads) and\n wants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.\n- Useful for adding the SMILES-reconstruction contrastive objective to a\n pretrained encoder without restarting pretraining from scratch (which\n `kermt-pretrain-scratch` would do at days-scale).\n\nFor continuing an existing hybrid or cmim ckpt: use `kermt-continue-pretrain`\ndirectly. For training a fresh model on a custom corpus: use\n`kermt-pretrain-scratch`.\n\n## Inputs\n\nRequired:\n\n- `--ckpt <path>` — grover_base ckpt to upgrade. Validated via\n `check_checkpoint.py --mode upgrade_to_hybrid`; rejected if the ckpt\n already has a contrast head or task FFN.\n- `--csv <path>` — pretrain corpus CSV. Same shape as\n `kermt-continue-pretrain`'s `--csv` input.\n\nOptional (same as `kermt-continue-pretrain`):\n\n- `--val-csv <path>` — separate validation CSV. Without it, prepare_data\n auto-splits by `--val-frac 0.1`.\n- Training-hyperparameter overrides (`--epochs N`, `--batch-size N`, lr triple,\n `--warmup-epochs F`, etc.).\n- `--vocab-loss-weight F` / `--latent-dim N` / `--contrastive-temperature F`.\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n logging (run name honored only alongside a project). Off by default.\n- `--gpus 0,2`.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout.\n\n1. **Pre-flight: check_system** (same as `kermt-continue-pretrain` step 1).\n\n2. **Compute run directory:**\n ```\n RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n ```\n\n3. **Validate the input ckpt with `check_checkpoint --mode upgrade_to_hybrid`.**\n Abort on `ok: false`. The validator rejects ckpts that already have\n contrast head (suggest `kermt-continue-pretrain`) or task FFN heads\n (the ckpt has been finetuned; suggest using the original pretrain\n checkpoint).\n\n4. **Validate the corpus** via `check_data --mode pretrain`. Abort on\n `ok: false`.\n\n5. **Prepare the data** with `--mode pretrain` — *without* `--vocab-dir`.\n The upgrade builds fresh vocab heads sized to the corpus's vocab, so we\n want `prepare_data` to produce a new vocab from the corpus rather than\n passing through the ckpt's old vocab (which may not even exist for\n encoder-only legacy grover_base ckpts):\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n --csv /data/<basename> --out /runs/data \\\\\n [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n ```\n The output manifest has `vocab_source: \"built_fresh\"` and includes a\n `smiles_vocab` (built from the corpus, needed for the new decoder).\n\n6. **Upgrade the ckpt.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\n \"python agent/scripts/upgrade_to_hybrid.py \\\\\n --ckpt /ckpt \\\\\n --prepare-manifest /runs/data/prepare_data.json \\\\\n --out /runs/upgraded.pt\"\n ```\n Surface the JSON summary to the user — especially `warnings[]`, which\n includes any encoder-arch drift notes (e.g. legacy GROVER had two extra\n `act_func_*` keys that modern KERMTEmbedding doesn't) and the\n pretrain_ddp.py `--backbone` argparse-restriction note if the upgraded\n ckpt's backbone is anything other than `gtrans`.\n\n7. **Estimate runtime + confirm with the user.** Same heuristic as\n `kermt-continue-pretrain` (corpus size × epochs × GPU count → wall time).\n\n8. **Launch the runner detached.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n --name kermt-add-cmim-pretrain-<ts> \\\\\n --run-dir $RUN_DIR -- \\\\\n \"python agent/scripts/run_pretrain_local.py \\\\\n --ckpt /runs/upgraded.pt \\\\\n --prepare-manifest /runs/data/prepare_data.json \\\\\n --out /runs \\\\\n [--epochs N --batch-size N ...]\"\n ```\n The runner sees the upgraded ckpt as `model_type: hybrid`, so it auto-dispatches\n `--pretrain_mode hybrid --vocab_loss_weight 1.0` with smiles_vocab plumbed\n through.\n\n9. **Report to the user** with the upgraded ckpt path + the same run.json\n pointer / log path / tensorboard URL pattern as `kermt-continue-pretrain`.\n\n## Hard rules\n\n- **Never modify the user's input ckpt.** The upgrade writes a new file at\n `<run_dir>/upgraded.pt`; the source ckpt stays untouched.\n- **Vocab heads are always fresh.** Even if the input grover_base has vocab\n heads, they're discarded and rebuilt sized to the new corpus's vocab.\n Continue-pretraining the upgraded ckpt will train those new heads alongside\n the decoder.\n- **Don't auto-relax `--backbone` choices.** If the upgrade warning fires\n because the input ckpt's backbone isn't `gtrans` (e.g. legacy `dualtrans`),\n surface the warning and ask the user. Do NOT silently modify parsing.py to\n add the legacy backbone to the choices list.\n\n## Common errors\n\n- `check_checkpoint rejected the ckpt` with model_type=hybrid or cmim →\n user's ckpt already has a contrast head. Redirect to\n `kermt-continue-pretrain`.\n- `check_checkpoint rejected the ckpt` with task_ffn=true → the ckpt has\n been finetuned. The upgrade workflow only supports pretrain checkpoints.\n- `prepare manifest missing smiles_vocab` → prepare_data was invoked with\n `--skip-vocab` or some equivalent that omitted the smiles vocab. Re-run\n prepare without those flags.\n- `unexpected key(s) in encoder load` warning → legacy GROVER architectures\n saved a couple of `act_func_*` weights that modern KERMTEmbedding doesn't\n use. Benign; the rest of the encoder loaded correctly.\n\n## What's in `run.json` after a successful run\n\nSame reproducibility fields as `kermt-continue-pretrain`, plus the upgrade step's\n`summary.json` is captured under the `inputs.upgrade_summary` path so the\nprovenance of the upgraded ckpt is auditable.\n\n## Replayability\n\nSame as `kermt-continue-pretrain`: `cmd_replay` rebuilds the\n`run_pretrain_local.py --ckpt <upgraded.pt> ...` invocation. To redo the\nfull add-cmim flow end-to-end, the user also needs the input grover_base\nckpt and the corpus — both are captured in the prepare_data and upgrade\nmanifests by absolute path.\n"
}SHA-256: ac8bcfa0c28ba2122c76e7ec32d737b401e3311f86d997ead58437cd3a6d74f9