{"id":17587,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.432Z","digest":"ac8bcfa0c28ba2122c76e7ec32d737b401e3311f86d997ead58437cd3a6d74f9","against":null,"payload":{"name":"kermt-add-cmim-pretrain","description":"Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.","included_files":[],"skill_md_contents":"---\nname: kermt-add-cmim-pretrain\ndescription: Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: workflow-skill\n  risk_tier: skill\n# Line/token budget: ~165 lines, ~1900 tokens — well within the\n# 500-line / 5000-token cap for skill files.\n---\n\n# kermt-add-cmim-pretrain\n\nConvert a grover_base checkpoint (legacy original-GROVER `grover.encoders.*`\nor modern `kermt.encoders.*`, with or without vocab heads) into a fully-formed\nhybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's\ncorpus as hybrid.\n\nThis is a thin wrapper: `upgrade_to_hybrid.py` produces a new ckpt that\nclassifies as `model_type: hybrid` via `check_checkpoint.py`, and the rest of\nthe workflow is identical to `kermt-continue-pretrain`.\n\n> **Status: experimental.** This workflow is functional end-to-end but has not\n> been benchmarked against the manuscript's from-scratch hybrid training (which\n> produces the released checkpoint). Use as an experimental alternative to\n> `kermt-pretrain-scratch` when you want to extend an existing grover_base\n> checkpoint rather than restart from random init. Validate downstream\n> performance on your own benchmark before relying on the upgraded ckpt for\n> production work.\n\n## Hardware requirements\n\nSame as `kermt-continue-pretrain` (the cMIM decoder adds parameters but not\nsubstantially; VRAM headroom should be fine). The upgrade step itself is\nfast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes\nGPU.\n\n## When to invoke\n\n- User has a grover_base checkpoint (encoder-only or with vocab heads) and\n  wants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.\n- Useful for adding the SMILES-reconstruction contrastive objective to a\n  pretrained encoder without restarting pretraining from scratch (which\n  `kermt-pretrain-scratch` would do at days-scale).\n\nFor continuing an existing hybrid or cmim ckpt: use `kermt-continue-pretrain`\ndirectly. For training a fresh model on a custom corpus: use\n`kermt-pretrain-scratch`.\n\n## Inputs\n\nRequired:\n\n- `--ckpt <path>` — grover_base ckpt to upgrade. Validated via\n  `check_checkpoint.py --mode upgrade_to_hybrid`; rejected if the ckpt\n  already has a contrast head or task FFN.\n- `--csv <path>` — pretrain corpus CSV. Same shape as\n  `kermt-continue-pretrain`'s `--csv` input.\n\nOptional (same as `kermt-continue-pretrain`):\n\n- `--val-csv <path>` — separate validation CSV. Without it, prepare_data\n  auto-splits by `--val-frac 0.1`.\n- Training-hyperparameter overrides (`--epochs N`, `--batch-size N`, lr triple,\n  `--warmup-epochs F`, etc.).\n- `--vocab-loss-weight F` / `--latent-dim N` / `--contrastive-temperature F`.\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n  logging (run name honored only alongside a project). Off by default.\n- `--gpus 0,2`.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout.\n\n1. **Pre-flight: check_system** (same as `kermt-continue-pretrain` step 1).\n\n2. **Compute run directory:**\n   ```\n   RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n   ```\n\n3. **Validate the input ckpt with `check_checkpoint --mode upgrade_to_hybrid`.**\n   Abort on `ok: false`. The validator rejects ckpts that already have\n   contrast head (suggest `kermt-continue-pretrain`) or task FFN heads\n   (the ckpt has been finetuned; suggest using the original pretrain\n   checkpoint).\n\n4. **Validate the corpus** via `check_data --mode pretrain`. Abort on\n   `ok: false`.\n\n5. **Prepare the data** with `--mode pretrain` — *without* `--vocab-dir`.\n   The upgrade builds fresh vocab heads sized to the corpus's vocab, so we\n   want `prepare_data` to produce a new vocab from the corpus rather than\n   passing through the ckpt's old vocab (which may not even exist for\n   encoder-only legacy grover_base ckpts):\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n            --csv /data/<basename> --out /runs/data \\\\\n            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n   ```\n   The output manifest has `vocab_source: \"built_fresh\"` and includes a\n   `smiles_vocab` (built from the corpus, needed for the new decoder).\n\n6. **Upgrade the ckpt.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/upgrade_to_hybrid.py \\\\\n            --ckpt /ckpt \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --out /runs/upgraded.pt\"\n   ```\n   Surface the JSON summary to the user — especially `warnings[]`, which\n   includes any encoder-arch drift notes (e.g. legacy GROVER had two extra\n   `act_func_*` keys that modern KERMTEmbedding doesn't) and the\n   pretrain_ddp.py `--backbone` argparse-restriction note if the upgraded\n   ckpt's backbone is anything other than `gtrans`.\n\n7. **Estimate runtime + confirm with the user.** Same heuristic as\n   `kermt-continue-pretrain` (corpus size × epochs × GPU count → wall time).\n\n8. **Launch the runner detached.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n       --name kermt-add-cmim-pretrain-<ts> \\\\\n       --run-dir $RUN_DIR -- \\\\\n       \"python agent/scripts/run_pretrain_local.py \\\\\n            --ckpt /runs/upgraded.pt \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --out /runs \\\\\n            [--epochs N --batch-size N ...]\"\n   ```\n   The runner sees the upgraded ckpt as `model_type: hybrid`, so it auto-dispatches\n   `--pretrain_mode hybrid --vocab_loss_weight 1.0` with smiles_vocab plumbed\n   through.\n\n9. **Report to the user** with the upgraded ckpt path + the same run.json\n   pointer / log path / tensorboard URL pattern as `kermt-continue-pretrain`.\n\n## Hard rules\n\n- **Never modify the user's input ckpt.** The upgrade writes a new file at\n  `<run_dir>/upgraded.pt`; the source ckpt stays untouched.\n- **Vocab heads are always fresh.** Even if the input grover_base has vocab\n  heads, they're discarded and rebuilt sized to the new corpus's vocab.\n  Continue-pretraining the upgraded ckpt will train those new heads alongside\n  the decoder.\n- **Don't auto-relax `--backbone` choices.** If the upgrade warning fires\n  because the input ckpt's backbone isn't `gtrans` (e.g. legacy `dualtrans`),\n  surface the warning and ask the user. Do NOT silently modify parsing.py to\n  add the legacy backbone to the choices list.\n\n## Common errors\n\n- `check_checkpoint rejected the ckpt` with model_type=hybrid or cmim →\n  user's ckpt already has a contrast head. Redirect to\n  `kermt-continue-pretrain`.\n- `check_checkpoint rejected the ckpt` with task_ffn=true → the ckpt has\n  been finetuned. The upgrade workflow only supports pretrain checkpoints.\n- `prepare manifest missing smiles_vocab` → prepare_data was invoked with\n  `--skip-vocab` or some equivalent that omitted the smiles vocab. Re-run\n  prepare without those flags.\n- `unexpected key(s) in encoder load` warning → legacy GROVER architectures\n  saved a couple of `act_func_*` weights that modern KERMTEmbedding doesn't\n  use. Benign; the rest of the encoder loaded correctly.\n\n## What's in `run.json` after a successful run\n\nSame reproducibility fields as `kermt-continue-pretrain`, plus the upgrade step's\n`summary.json` is captured under the `inputs.upgrade_summary` path so the\nprovenance of the upgraded ckpt is auditable.\n\n## Replayability\n\nSame as `kermt-continue-pretrain`: `cmd_replay` rebuilds the\n`run_pretrain_local.py --ckpt <upgraded.pt> ...` invocation. To redo the\nfull add-cmim flow end-to-end, the user also needs the input grover_base\nckpt and the corpus — both are captured in the prepare_data and upgrade\nmanifests by absolute path.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}