← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "kermt-pretrain-scratch",
"description": "Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.",
"included_files": [],
"skill_md_contents": "---\nname: kermt-pretrain-scratch\ndescription: Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n owner: evax@nvidia.com\n classification: workflow-skill\n risk_tier: skill\n# Line/token budget: ~210 lines, ~2400 tokens — well within the\n# 500-line / 5000-token cap for skill files. Most of the orchestration is\n# shared with kermt-continue-pretrain; the differences are documented below.\n---\n\n# kermt-pretrain-scratch\n\nPretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful\nwhen you want to retrain a model on a custom chemistry domain rather than\nextending one of the released checkpoints. **Significantly more expensive than\n`kermt-continue-pretrain`** — no warm start, so the loss curves need to descend\nfrom scratch over many epochs.\n\n## Hardware requirements\n\nSame as `kermt-continue-pretrain`:\n\n- **GPUs**: 1–N CUDA-capable. The runner auto-detects via\n `torch.cuda.device_count()`; `--gpus 0,2` overrides. Single-GPU fallback:\n `--batch_size 32 --save_interval 500`. Multi-GPU keeps defaults\n (`--batch_size 256` etc.). Note: `--gpus N` uses **torch.cuda** indexing,\n which can differ from `nvidia-smi`'s display order on multi-GPU hosts\n (PCI bus vs. CUDA enumeration). To target a specific physical GPU, set\n `CUDA_VISIBLE_DEVICES` before invoking, or run\n `python -c \"import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])\"`\n to confirm which device you're picking.\n- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware\n (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:\n\n | GPU class | VRAM | Suggested `--batch-size` |\n |--------------------------|------------|--------------------------|\n | L4, T4, V100 16 GB | 16–24 GB | 32–64 |\n | A100 40 GB, L40, A40 | 40–48 GB | 128 |\n | A100 80 GB, H100, H200 | 80 GB | 256 (default) |\n\n These are rough starting points — pass `--batch-size N` to override.\n- **Disk**: tens of GB for shards + vocab + checkpoints, scaled by epochs.\n- **Wall time**: this is the big difference. Pretraining from scratch on an\n 11M-mol corpus at 100 epochs typically takes **days even on a multi-GPU box**.\n The skill prints an estimate before launching; confirm with the user.\n\n## When to invoke\n\n- User wants to train a new model on a custom corpus (e.g. domain-specific\n chemistry that the released ckpts don't cover).\n- User wants to reproduce a pretrain config end-to-end without depending on a\n released ckpt.\n\nFor continuing an existing released ckpt, use `kermt-continue-pretrain`. For\nadding a cMIM decoder to an encoder-only grover_base ckpt, use\n`kermt-add-cmim-pretrain`.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — the pretrain corpus CSV with a `smiles` column. Single file\n by convention; multi-file corpora deferred. Use `--val-csv` for a separate\n validation set.\n- `--pretrain-target-mode {vocab|cmim|hybrid}` — which pretrain objective to\n use. **No default** — must be set explicitly so the user makes an informed\n choice:\n - `vocab` — original GROVER-style atom + bond vocab prediction (encoder-only\n output, lightweight).\n - `cmim` — contrastive + SMILES reconstruction objective. Requires building\n a SMILES vocab from the corpus.\n - `hybrid` — both vocab and contrastive objectives jointly (the\n state-of-the-art config from the KERMT manuscript).\n\nOptional:\n\n- `--val-csv <path>` — separate validation CSV. Without it, prepare_data\n auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).\n- Training-hyperparameter overrides: `--epochs N` / `--batch-size N` /\n `--init-lr F` / `--max-lr F` / `--final-lr F` / `--warmup-epochs F` /\n `--weight-decay F` / `--dropout F` / `--save-interval N` / `--seed N`.\n Anything not given is filled from `agent/config/defaults_pretrain.json`.\n- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /\n `--contrastive-temperature F` (cmim and hybrid only).\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run\n name is honored only alongside a project. Off by default.\n- `--gpus 0,2` — restrict to a GPU subset.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout.\n\n1. **Pre-flight: ensure container + system probe** (same as\n `kermt-continue-pretrain` step 1). Refuse to proceed if `check_system`\n reports gaps.\n\n2. **Compute run directory.**\n ```\n RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n ```\n\n3. **Validate the corpus** (no ckpt to validate, so this is the only input\n check):\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n \"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>\"\n ```\n Abort on `ok: false`.\n\n4. **Prepare the data** — no vocab pass-through (we want fresh vocab from\n corpus):\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n --csv /data/<basename> --out /runs/data \\\\\n [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n ```\n Outputs land at `$RUN_DIR/data/prepare_data.json` with\n `vocab_source: \"built_fresh\"`.\n\n5. **Estimate runtime + warn loudly.** This is critical for pretrain-from-scratch:\n - \"Pretraining from scratch is days-scale even on multi-GPU; the released\n KERMT checkpoints were each trained on millions of molecules for hundreds\n of GPU-hours. If you mainly want to leverage existing knowledge for a\n downstream task, consider `kermt-continue-pretrain` from a released ckpt\n instead, which converges in hours instead of days.\"\n - Show the corpus size × epochs × GPU count → estimated wall time.\n - Ask for explicit confirmation unless `--yes` was given.\n\n6. **Launch the runner detached.**\n ```\n $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n --name kermt-pretrain-scratch-<ts> \\\\\n --run-dir $RUN_DIR -- \\\\\n \"python agent/scripts/run_pretrain_local.py \\\\\n --from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\\\\n --prepare-manifest /runs/data/prepare_data.json \\\\\n --out /runs \\\\\n [--epochs N --batch-size N ...]\"\n ```\n Note: NO `--ckpt` flag (the runner refuses if both `--from-scratch` and\n `--ckpt` are given). The runner uses the `arch` group from\n `agent/config/defaults_pretrain.json` to size the model.\n\n7. **Report to the user.** Always include all of the following — do not\n omit the TensorBoard line under output-length pressure:\n - Container name + id\n - `$RUN_DIR/run.json` (the manifest with `workflow: pretrain-scratch`,\n `from_scratch: true`, `vocab_check: null`, `arch` from defaults, full\n `cmd_replay`)\n - Log file: `$RUN_DIR/logs/pretrain_ddp.log`\n - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir\n $RUN_DIR/logs/tb`)\n - Suggest `kermt-monitor <RUN_DIR>` for progress.\n\n## Hard rules\n\n- **Never accept a `--ckpt` flag.** From-scratch is exclusive with input\n ckpt — the runner enforces this; the skill should too.\n- **Never silently default `--pretrain-target-mode`.** This is a significant\n architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user\n if not given on the CLI.\n- **Strong warning before launching.** From-scratch pretrain is the most\n expensive workflow. The user needs to know what they're committing to.\n\n## Common errors\n\n- `--pretrain-target-mode is required when --from-scratch is set` → user\n forgot the mode flag. Prompt.\n- `--from-scratch is incompatible with --ckpt` → user provided both; ask which\n one they meant.\n- `defaults_pretrain.json has no arch group` → repo state issue (should never\n happen on a fresh clone); points the user at running `kermt-setup` again.\n\n## What's in the manifest after a from-scratch run\n\nSame reproducibility fields as continue-pretrain (`repo.commit`, `kermt_image`,\n`cmd_replay`, `args_applied`), plus:\n\n- `workflow`: `\"pretrain-scratch\"`\n- `from_scratch`: `true`\n- `inputs.ckpt`: `null`\n- `ckpt_symlink`: `null`\n- `vocab_check`: `null` (not verified — vocab built from corpus is\n authoritative for from-scratch)\n- `arch`: the values pulled from `agent/config/defaults_pretrain.json`'s\n `arch` group (with any future CLI overrides applied).\n\n## Replayability\n\nSame as continue-pretrain: `cmd_replay` is a copy-pasteable command. If\n`ok_to_replay: false`, the kermt repo working tree was dirty at launch\ntime — check `repo.commit` and `git checkout` it first.\n"
}SHA-256: 44f815427cd2f8620a8f4b51a59819c9a6e1f947d7e0c06295e6742169d8c635