← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "kermt-continue-pretrain",
"description": "Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).",
"included_files": [],
"skill_md_contents": "---\nname: kermt-continue-pretrain\ndescription: Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n owner: evax@nvidia.com\n classification: workflow-skill\n risk_tier: skill\n# Line/token budget: this file is targeted at ~250 lines / ~3000 tokens —\n# well within the 500-line / 5000-token cap. Long examples live in\n# agent/scripts/run_pretrain_local.py's docstring.\n---\n\n# kermt-continue-pretrain\n\nContinue pretraining from a user-supplied KERMT checkpoint (grover_base /\ncmim / hybrid). The skill is the workflow orchestrator: it validates inputs,\nprepares the corpus, launches the runner, and returns a run directory.\n\n## Hardware requirements\n\n- **GPUs**: 1–N CUDA-capable NVIDIA GPUs. The runner auto-detects via\n `torch.cuda.device_count()`; `--gpus 0,2` overrides. On a single GPU the\n runner falls back to `--batch_size 32 --save_interval 500`; on multi-GPU\n it uses the `defaults_pretrain.json` values (currently `batch_size 256`).\n Note: `--gpus N` uses **torch.cuda** indexing, which can differ from\n `nvidia-smi`'s display order on multi-GPU hosts (PCI bus vs. CUDA\n enumeration). To target a specific physical GPU, set `CUDA_VISIBLE_DEVICES`\n before invoking, or run\n `python -c \"import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])\"`\n to confirm which device you're picking.\n- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware\n (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:\n\n | GPU class | VRAM | Suggested `--batch-size` |\n |--------------------------|------------|--------------------------|\n | L4, T4, V100 16 GB | 16–24 GB | 32–64 |\n | A100 40 GB, L40, A40 | 40–48 GB | 128 |\n | A100 80 GB, H100, H200 | 80 GB | 256 (default) |\n\n These are rough starting points — pass `--batch-size N` to override.\n- **Disk**: tens of GB depending on corpus size + epochs (each checkpoint\n is several hundred MB).\n- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).\n `kermt-setup` validates this up-front.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — the pretrain CSV (single column `smiles`). If you have\n separate train/val CSVs, pass `--val-csv <path>` too.\n\nCheckpoint (optional — defaults to the released model if omitted):\n\n- `--ckpt <path>` — the input pretrain checkpoint to continue from. Must be\n a grover_base (with vocab heads), cmim, or hybrid ckpt; the validator\n rejects everything else with a redirect to the correct workflow. **If\n omitted**, the skill offers to download the released pretrained hybrid model\n **nvidia/NV-KERMT-70M-v2** and continue-pretrain from it — see \"Resolve &\n validate the checkpoint\" (workflow step 3). The released bundle ships its\n three vocab files alongside the ckpt, so the authoritative-vocab pass-through\n (step 5) works automatically.\n- `--pretrained-release` — explicit opt-in to use the released model without\n the interactive prompt (for non-interactive / agent runs). Mutually\n exclusive with `--ckpt`.\n- `--model-dir <dir>` — where to save the downloaded bundle (default\n `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is\n reused, not re-downloaded.\n\nOptional:\n\n- `--val-csv <path>` — separate validation CSV. Without it, the prep step\n auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).\n- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /\n `--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /\n `--save-interval N` / `--seed N` — training-hyperparameter overrides.\n Anything not given is filled from `agent/config/defaults_pretrain.json`.\n- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /\n `--contrastive-temperature F` (cmim and hybrid only) — loss / decoder\n overrides.\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run\n name is honored only alongside a project. Off by default. (Independent of the\n ckpt's `wandb_run_id` continuity handling under `--resume`.)\n- `--resume` — see \"Modes\" section below.\n- `--gpus 0,2` — restrict to a GPU subset. Default uses all visible GPUs.\n- `--from-prepare <dir>` — skip the prepare step and reuse an existing\n `prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.\n\n## Modes\n\nThe runner has two modes for ingesting the input ckpt, dispatched on whether\n`--resume` is set. Pick based on intent:\n\n### Default (fresh-schedule continue-pretrain)\n\n**Use when**: you have a finished pretrain ckpt and want to continue training\nit — on a new corpus, with a different objective, or just for more epochs\nthan its original plan. The previous training's step counter and schedule\nshape are no longer relevant; you want a new learning-rate schedule for the\nnew run.\n\n**What gets loaded from the ckpt**:\n- ✓ Model weights (encoder + vocab heads + contrast head + decoder, whatever\n is there)\n- ✓ Optimizer state (Adam's running m1/m2 moments — warm-starts the new\n schedule so the first few hundred steps aren't dominated by noisy\n gradient-estimate startup)\n- ✗ Scheduler step counter (reset to 0)\n- ✗ Epoch counter (reset to 0)\n- ✗ Batch counter (reset to 0)\n- ✗ wandb run id (new wandb run, not a continuation)\n\n**Schedule shape** (init/max/final LR, warmup epochs, total epochs): from\nyour CLI args or `defaults_pretrain.json`. A fresh NoamLR is constructed\nfrom these values and starts at step 0.\n\n### `--resume` (true resume)\n\n**Use when**: a previous run was interrupted (crash, OOM, Ctrl-C) and you\nwant to pick up exactly where it left off — same dataset, same schedule,\nsame training trajectory.\n\n**What gets loaded from the ckpt**: **everything** in the\n`save_model_for_restart` format. Model weights + optimizer state +\nscheduler_step + epoch + batch_idx + wandb_run_id are all restored. The\nnew run continues from the saved step in the saved schedule (which is\nrecovered from the ckpt's `saved_args`). Mid-epoch resume works too —\n`pretrain_ddp.py`'s sampler skip-count picks up at the saved batch index\nwithin the saved epoch.\n\n**Schedule shape**: inherited from the ckpt's `saved_args`. CLI overrides\nof any schedule flag (`--epochs / --warmup-epochs / --init-lr / --max-lr /\n--final-lr`) are **rejected with a hard error** — pure resume means pure\nresume; if you want to change the schedule, drop `--resume` and start a\nfresh-schedule run.\n\n**Requirements**: the ckpt must have been saved via `save_model_for_restart`\n(i.e., carry `optimizer / scheduler_step / epoch / batch_idx` keys). If\nany of these is missing, the runner errors with a clear message and\nsuggests dropping `--resume`.\n\nThe default mode is the right choice ~90% of the time. Reach for `--resume`\nonly when you genuinely need to continue a single interrupted training\nrun.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout, and assume\n`kermt-setup` has already built `kermt:latest`. All paths below are on the\nhost; the helper bind-mounts them at known container paths.\n\n1. **Pre-flight: ensure container + system probe.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c \"\n import json, sys; d = json.load(sys.stdin)\n if not d['ok']:\n print('System check failed:', d['gaps']); sys.exit(1)\n print(f'OK: {len(d[\\\"gpus\\\"])} GPU(s); {d[\\\"disk\\\"][\\\"free_gb\\\"]} GB free; CUDA via container toolkit')\n \"\n ```\n Surface any `gaps` to the user. Refuse to proceed if `ok: false`.\n\n2. **Compute run directory.**\n ```\n RUN_DIR=$KERMT_REPO/runs/continue-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n ```\n\n3. **Resolve & validate the checkpoint.**\n\n **Resolve — only if `--ckpt` was omitted.** Default to the released\n pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:\n - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:\n \"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2\n (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)\n and continue-pretrain from it? [y/N]\". **Never download without an\n explicit yes** (or `--pretrained-release`). If both `--ckpt` and\n `--pretrained-release` are given, abort — they conflict.\n - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor\n `--model-dir <dir>` if given. An already-complete bundle is reused.\n - **Download** (foreground; ~282 MB on first fetch):\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \\\n \"python agent/scripts/fetch_released_model.py --out /model\"\n ```\n Parse the JSON; abort on `ok: false` (surface `errors`). On success set\n `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`. The bundle's three\n vocab files land in `<save-dir>` too, so step 5's `--vocab-dir`\n auto-detection (which looks in the ckpt's parent directory) finds them\n with no extra work.\n\n **Validate** the resolved (or user-provided) ckpt:\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \\\n \"python agent/scripts/check_checkpoint.py --mode continue_pretrain --ckpt /ckpt\"\n ```\n Parse the JSON. Abort on `ok: false`, showing the error verbatim. The error\n message redirects the user to `kermt-add-cmim-pretrain` for encoder-only\n ckpts, or to `kermt-finetune` for finetuned ckpts.\n\n4. **Validate the data.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n \"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>\"\n ```\n Abort on `ok: false`.\n\n5. **Prepare the data** (skip if `--from-prepare` given).\n **Pass the ckpt's vocab through.** Look in the ckpt's parent directory for\n the conventional `pretrain_atom_vocab.{json,pkl}`, `pretrain_bond_vocab.{json,pkl}`,\n and `pretrain_smiles_vocab.pkl` files (the bundling convention for released\n models; see `agent/README.md` \"Released models\" section). If all three are\n present, auto-pass via `--vocab-dir <ckpt_parent_dir>`. If only some are\n present, pass them via explicit flags (`--atom-vocab`, `--bond-vocab`,\n `--smiles-vocab`). If none are present, ask the user for `--vocab-dir` — or\n refuse to proceed, because rebuilding a fresh vocab from the new corpus\n would silently mismatch the ckpt's vocab heads (the ckpt's vocab is\n authoritative for continue-pretrain).\n\n Note the **two-layer mount pattern**: pass the host directory to\n `kermt_container.sh --vocab-dir` (which mounts it at `/vocab` inside the\n container), and reference `/vocab` from the inner `prepare_data.py`\n command. The same pattern applies to every host path the inner command\n needs to read (`--data <host-csv>` → `/data/<basename>`,\n `--ckpt <host-ckpt>` → `/ckpt`).\n\n ```\n VOCAB_DIR=$(dirname <user-ckpt>)\n $KERMT_REPO/agent/scripts/kermt_container.sh run \\\n --data <user-csv> --vocab-dir $VOCAB_DIR --run-dir $RUN_DIR -- \\\n \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n --csv /data/<basename> --out /runs/data \\\\\n --vocab-dir /vocab \\\\\n [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n ```\n Outputs land at `$RUN_DIR/data/prepare_data.json` with\n `vocab_source: \"user_provided\"`. The runner step 7 will verify the vocab\n files' entry counts match the ckpt's vocab-head sizes and refuse to launch\n on mismatch.\n\n6. **Estimate runtime + confirm with user.**\n - Pretrain wall time depends on corpus size × epochs × GPU count.\n - Tell the user the estimate; ask \"proceed?\" unless `--yes` flag was given\n (agent-non-interactive case).\n - Example estimate template:\n `~N hours on K GPUs for E epochs over M molecules (~steps/epoch × seconds/step)`.\n\n7. **Launch the runner detached.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n --name kermt-continue-pretrain-<ts> \\\\\n --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\\\n \"python agent/scripts/run_pretrain_local.py \\\\\n --ckpt /ckpt \\\\\n --prepare-manifest /runs/data/prepare_data.json \\\\\n --out /runs \\\\\n [--epochs N --batch-size N --init-lr F ...]\"\n ```\n Returns the container name + id + log file path.\n\n8. **Report to the user.** Output a short summary:\n - Container name + id\n - `$RUN_DIR/run.json` (the manifest with cmd_replay + image digest)\n - Log file: `$RUN_DIR/logs/pretrain_ddp.log`\n - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir\n $RUN_DIR/logs/tb`)\n - Suggest invoking `kermt-monitor <RUN_DIR>` to check progress.\n\n## Hard rules\n\n- **Never download the released model without consent.** When `--ckpt` is\n omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user \"yes\"\n or an explicit `--pretrained-release` flag. `--ckpt` and\n `--pretrained-release` are mutually exclusive.\n- **Never modify the user's input ckpt.** The runner symlinks it into the\n save_dir; the symlink is what pretrain_ddp.py auto-resumes from. The\n source file stays untouched.\n- **Never silently override arch.** If the user passes a `--hidden-size`\n etc. that doesn't match the ckpt-derived value, the runner aborts loudly.\n Arch params come from the ckpt, period.\n- **Never block on the long-running pretrain itself.** The runner is invoked\n via `run_detached`; the skill returns immediately after step 8. Use\n `kermt-monitor` for progress.\n- **Echo applied defaults back to the user.** The `args_applied` field of\n `run.json` records every flag's value + source (user / default-config /\n auto-1gpu / auto-multi-gpu). Skill should surface a summary of any flag\n not user-specified so the user knows what was assumed.\n\n## Common errors\n\n- `model_type='finetuned'` rejected → the ckpt is a downstream finetune,\n not a pretrain. The error redirects to the relevant workflow.\n- `grover_base ckpt has no vocab head` → encoder-only ckpt (e.g. the\n original-grover `grover_base.pt`). The error redirects to\n `kermt-add-cmim-pretrain`.\n- `prepare_data manifest is missing required outputs` → user passed\n `--from-prepare` to a directory where prepare was run with `--skip-vocab`\n or `--skip-split`. Re-run prepare without those flags.\n- `--gpus all` not available → install `nvidia-container-toolkit`; check\n `kermt_container.sh check_system`.\n\n## Replayability\n\nThe `run.json` `cmd_replay` field is a single-line command that re-runs the\npretrain with the same inputs, hyperparameters, and arch. To replay:\n\n```bash\n# Inside the kermt container:\n$(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n\nIf `ok_to_replay: false` in the manifest (because the kermt repo working\ntree was dirty at launch time), the replay may not be bit-exact — pin the\nexact commit via the `repo.commit` field and `git checkout` it\nfirst.\n"
}SHA-256: d8684b0dc4a99478026f8242a09251fa6a1dd42fdf0dc07aa2fdfcb3b0a33e1a