{"id":17588,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.470Z","digest":"d8684b0dc4a99478026f8242a09251fa6a1dd42fdf0dc07aa2fdfcb3b0a33e1a","against":null,"payload":{"name":"kermt-continue-pretrain","description":"Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).","included_files":[],"skill_md_contents":"---\nname: kermt-continue-pretrain\ndescription: Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: workflow-skill\n  risk_tier: skill\n# Line/token budget: this file is targeted at ~250 lines / ~3000 tokens —\n# well within the 500-line / 5000-token cap. Long examples live in\n# agent/scripts/run_pretrain_local.py's docstring.\n---\n\n# kermt-continue-pretrain\n\nContinue pretraining from a user-supplied KERMT checkpoint (grover_base /\ncmim / hybrid). The skill is the workflow orchestrator: it validates inputs,\nprepares the corpus, launches the runner, and returns a run directory.\n\n## Hardware requirements\n\n- **GPUs**: 1–N CUDA-capable NVIDIA GPUs. The runner auto-detects via\n  `torch.cuda.device_count()`; `--gpus 0,2` overrides. On a single GPU the\n  runner falls back to `--batch_size 32 --save_interval 500`; on multi-GPU\n  it uses the `defaults_pretrain.json` values (currently `batch_size 256`).\n  Note: `--gpus N` uses **torch.cuda** indexing, which can differ from\n  `nvidia-smi`'s display order on multi-GPU hosts (PCI bus vs. CUDA\n  enumeration). To target a specific physical GPU, set `CUDA_VISIBLE_DEVICES`\n  before invoking, or run\n  `python -c \"import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])\"`\n  to confirm which device you're picking.\n- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware\n  (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:\n\n  | GPU class                | VRAM       | Suggested `--batch-size` |\n  |--------------------------|------------|--------------------------|\n  | L4, T4, V100 16 GB       | 16–24 GB   | 32–64                    |\n  | A100 40 GB, L40, A40     | 40–48 GB   | 128                      |\n  | A100 80 GB, H100, H200   | 80 GB      | 256 (default)            |\n\n  These are rough starting points — pass `--batch-size N` to override.\n- **Disk**: tens of GB depending on corpus size + epochs (each checkpoint\n  is several hundred MB).\n- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).\n  `kermt-setup` validates this up-front.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — the pretrain CSV (single column `smiles`). If you have\n  separate train/val CSVs, pass `--val-csv <path>` too.\n\nCheckpoint (optional — defaults to the released model if omitted):\n\n- `--ckpt <path>` — the input pretrain checkpoint to continue from. Must be\n  a grover_base (with vocab heads), cmim, or hybrid ckpt; the validator\n  rejects everything else with a redirect to the correct workflow. **If\n  omitted**, the skill offers to download the released pretrained hybrid model\n  **nvidia/NV-KERMT-70M-v2** and continue-pretrain from it — see \"Resolve &\n  validate the checkpoint\" (workflow step 3). The released bundle ships its\n  three vocab files alongside the ckpt, so the authoritative-vocab pass-through\n  (step 5) works automatically.\n- `--pretrained-release` — explicit opt-in to use the released model without\n  the interactive prompt (for non-interactive / agent runs). Mutually\n  exclusive with `--ckpt`.\n- `--model-dir <dir>` — where to save the downloaded bundle (default\n  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is\n  reused, not re-downloaded.\n\nOptional:\n\n- `--val-csv <path>` — separate validation CSV. Without it, the prep step\n  auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).\n- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /\n  `--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /\n  `--save-interval N` / `--seed N` — training-hyperparameter overrides.\n  Anything not given is filled from `agent/config/defaults_pretrain.json`.\n- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /\n  `--contrastive-temperature F` (cmim and hybrid only) — loss / decoder\n  overrides.\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n  logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run\n  name is honored only alongside a project. Off by default. (Independent of the\n  ckpt's `wandb_run_id` continuity handling under `--resume`.)\n- `--resume` — see \"Modes\" section below.\n- `--gpus 0,2` — restrict to a GPU subset. Default uses all visible GPUs.\n- `--from-prepare <dir>` — skip the prepare step and reuse an existing\n  `prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.\n\n## Modes\n\nThe runner has two modes for ingesting the input ckpt, dispatched on whether\n`--resume` is set. Pick based on intent:\n\n### Default (fresh-schedule continue-pretrain)\n\n**Use when**: you have a finished pretrain ckpt and want to continue training\nit — on a new corpus, with a different objective, or just for more epochs\nthan its original plan. The previous training's step counter and schedule\nshape are no longer relevant; you want a new learning-rate schedule for the\nnew run.\n\n**What gets loaded from the ckpt**:\n- ✓ Model weights (encoder + vocab heads + contrast head + decoder, whatever\n  is there)\n- ✓ Optimizer state (Adam's running m1/m2 moments — warm-starts the new\n  schedule so the first few hundred steps aren't dominated by noisy\n  gradient-estimate startup)\n- ✗ Scheduler step counter (reset to 0)\n- ✗ Epoch counter (reset to 0)\n- ✗ Batch counter (reset to 0)\n- ✗ wandb run id (new wandb run, not a continuation)\n\n**Schedule shape** (init/max/final LR, warmup epochs, total epochs): from\nyour CLI args or `defaults_pretrain.json`. A fresh NoamLR is constructed\nfrom these values and starts at step 0.\n\n### `--resume` (true resume)\n\n**Use when**: a previous run was interrupted (crash, OOM, Ctrl-C) and you\nwant to pick up exactly where it left off — same dataset, same schedule,\nsame training trajectory.\n\n**What gets loaded from the ckpt**: **everything** in the\n`save_model_for_restart` format. Model weights + optimizer state +\nscheduler_step + epoch + batch_idx + wandb_run_id are all restored. The\nnew run continues from the saved step in the saved schedule (which is\nrecovered from the ckpt's `saved_args`). Mid-epoch resume works too —\n`pretrain_ddp.py`'s sampler skip-count picks up at the saved batch index\nwithin the saved epoch.\n\n**Schedule shape**: inherited from the ckpt's `saved_args`. CLI overrides\nof any schedule flag (`--epochs / --warmup-epochs / --init-lr / --max-lr /\n--final-lr`) are **rejected with a hard error** — pure resume means pure\nresume; if you want to change the schedule, drop `--resume` and start a\nfresh-schedule run.\n\n**Requirements**: the ckpt must have been saved via `save_model_for_restart`\n(i.e., carry `optimizer / scheduler_step / epoch / batch_idx` keys). If\nany of these is missing, the runner errors with a clear message and\nsuggests dropping `--resume`.\n\nThe default mode is the right choice ~90% of the time. Reach for `--resume`\nonly when you genuinely need to continue a single interrupted training\nrun.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout, and assume\n`kermt-setup` has already built `kermt:latest`. All paths below are on the\nhost; the helper bind-mounts them at known container paths.\n\n1. **Pre-flight: ensure container + system probe.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c \"\n   import json, sys; d = json.load(sys.stdin)\n   if not d['ok']:\n       print('System check failed:', d['gaps']); sys.exit(1)\n   print(f'OK: {len(d[\\\"gpus\\\"])} GPU(s); {d[\\\"disk\\\"][\\\"free_gb\\\"]} GB free; CUDA via container toolkit')\n   \"\n   ```\n   Surface any `gaps` to the user. Refuse to proceed if `ok: false`.\n\n2. **Compute run directory.**\n   ```\n   RUN_DIR=$KERMT_REPO/runs/continue-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n   ```\n\n3. **Resolve & validate the checkpoint.**\n\n   **Resolve — only if `--ckpt` was omitted.** Default to the released\n   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:\n   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:\n     \"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2\n     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)\n     and continue-pretrain from it? [y/N]\". **Never download without an\n     explicit yes** (or `--pretrained-release`). If both `--ckpt` and\n     `--pretrained-release` are given, abort — they conflict.\n   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor\n     `--model-dir <dir>` if given. An already-complete bundle is reused.\n   - **Download** (foreground; ~282 MB on first fetch):\n     ```\n     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \\\n         \"python agent/scripts/fetch_released_model.py --out /model\"\n     ```\n     Parse the JSON; abort on `ok: false` (surface `errors`). On success set\n     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`. The bundle's three\n     vocab files land in `<save-dir>` too, so step 5's `--vocab-dir`\n     auto-detection (which looks in the ckpt's parent directory) finds them\n     with no extra work.\n\n   **Validate** the resolved (or user-provided) ckpt:\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \\\n       \"python agent/scripts/check_checkpoint.py --mode continue_pretrain --ckpt /ckpt\"\n   ```\n   Parse the JSON. Abort on `ok: false`, showing the error verbatim. The error\n   message redirects the user to `kermt-add-cmim-pretrain` for encoder-only\n   ckpts, or to `kermt-finetune` for finetuned ckpts.\n\n4. **Validate the data.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n       \"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>\"\n   ```\n   Abort on `ok: false`.\n\n5. **Prepare the data** (skip if `--from-prepare` given).\n   **Pass the ckpt's vocab through.** Look in the ckpt's parent directory for\n   the conventional `pretrain_atom_vocab.{json,pkl}`, `pretrain_bond_vocab.{json,pkl}`,\n   and `pretrain_smiles_vocab.pkl` files (the bundling convention for released\n   models; see `agent/README.md` \"Released models\" section). If all three are\n   present, auto-pass via `--vocab-dir <ckpt_parent_dir>`. If only some are\n   present, pass them via explicit flags (`--atom-vocab`, `--bond-vocab`,\n   `--smiles-vocab`). If none are present, ask the user for `--vocab-dir` — or\n   refuse to proceed, because rebuilding a fresh vocab from the new corpus\n   would silently mismatch the ckpt's vocab heads (the ckpt's vocab is\n   authoritative for continue-pretrain).\n\n   Note the **two-layer mount pattern**: pass the host directory to\n   `kermt_container.sh --vocab-dir` (which mounts it at `/vocab` inside the\n   container), and reference `/vocab` from the inner `prepare_data.py`\n   command. The same pattern applies to every host path the inner command\n   needs to read (`--data <host-csv>` → `/data/<basename>`,\n   `--ckpt <host-ckpt>` → `/ckpt`).\n\n   ```\n   VOCAB_DIR=$(dirname <user-ckpt>)\n   $KERMT_REPO/agent/scripts/kermt_container.sh run \\\n       --data <user-csv> --vocab-dir $VOCAB_DIR --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n            --csv /data/<basename> --out /runs/data \\\\\n            --vocab-dir /vocab \\\\\n            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n   ```\n   Outputs land at `$RUN_DIR/data/prepare_data.json` with\n   `vocab_source: \"user_provided\"`. The runner step 7 will verify the vocab\n   files' entry counts match the ckpt's vocab-head sizes and refuse to launch\n   on mismatch.\n\n6. **Estimate runtime + confirm with user.**\n   - Pretrain wall time depends on corpus size × epochs × GPU count.\n   - Tell the user the estimate; ask \"proceed?\" unless `--yes` flag was given\n     (agent-non-interactive case).\n   - Example estimate template:\n     `~N hours on K GPUs for E epochs over M molecules (~steps/epoch × seconds/step)`.\n\n7. **Launch the runner detached.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n       --name kermt-continue-pretrain-<ts> \\\\\n       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\\\n       \"python agent/scripts/run_pretrain_local.py \\\\\n            --ckpt /ckpt \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --out /runs \\\\\n            [--epochs N --batch-size N --init-lr F ...]\"\n   ```\n   Returns the container name + id + log file path.\n\n8. **Report to the user.** Output a short summary:\n   - Container name + id\n   - `$RUN_DIR/run.json` (the manifest with cmd_replay + image digest)\n   - Log file: `$RUN_DIR/logs/pretrain_ddp.log`\n   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir\n     $RUN_DIR/logs/tb`)\n   - Suggest invoking `kermt-monitor <RUN_DIR>` to check progress.\n\n## Hard rules\n\n- **Never download the released model without consent.** When `--ckpt` is\n  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user \"yes\"\n  or an explicit `--pretrained-release` flag. `--ckpt` and\n  `--pretrained-release` are mutually exclusive.\n- **Never modify the user's input ckpt.** The runner symlinks it into the\n  save_dir; the symlink is what pretrain_ddp.py auto-resumes from. The\n  source file stays untouched.\n- **Never silently override arch.** If the user passes a `--hidden-size`\n  etc. that doesn't match the ckpt-derived value, the runner aborts loudly.\n  Arch params come from the ckpt, period.\n- **Never block on the long-running pretrain itself.** The runner is invoked\n  via `run_detached`; the skill returns immediately after step 8. Use\n  `kermt-monitor` for progress.\n- **Echo applied defaults back to the user.** The `args_applied` field of\n  `run.json` records every flag's value + source (user / default-config /\n  auto-1gpu / auto-multi-gpu). Skill should surface a summary of any flag\n  not user-specified so the user knows what was assumed.\n\n## Common errors\n\n- `model_type='finetuned'` rejected → the ckpt is a downstream finetune,\n  not a pretrain. The error redirects to the relevant workflow.\n- `grover_base ckpt has no vocab head` → encoder-only ckpt (e.g. the\n  original-grover `grover_base.pt`). The error redirects to\n  `kermt-add-cmim-pretrain`.\n- `prepare_data manifest is missing required outputs` → user passed\n  `--from-prepare` to a directory where prepare was run with `--skip-vocab`\n  or `--skip-split`. Re-run prepare without those flags.\n- `--gpus all` not available → install `nvidia-container-toolkit`; check\n  `kermt_container.sh check_system`.\n\n## Replayability\n\nThe `run.json` `cmd_replay` field is a single-line command that re-runs the\npretrain with the same inputs, hyperparameters, and arch. To replay:\n\n```bash\n# Inside the kermt container:\n$(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n\nIf `ok_to_replay: false` in the manifest (because the kermt repo working\ntree was dirty at launch time), the replay may not be bit-exact — pin the\nexact commit via the `repo.commit` field and `git checkout` it\nfirst.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}