{"id":17595,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.640Z","digest":"44f815427cd2f8620a8f4b51a59819c9a6e1f947d7e0c06295e6742169d8c635","against":null,"payload":{"name":"kermt-pretrain-scratch","description":"Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.","included_files":[],"skill_md_contents":"---\nname: kermt-pretrain-scratch\ndescription: Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: workflow-skill\n  risk_tier: skill\n# Line/token budget: ~210 lines, ~2400 tokens — well within the\n# 500-line / 5000-token cap for skill files. Most of the orchestration is\n# shared with kermt-continue-pretrain; the differences are documented below.\n---\n\n# kermt-pretrain-scratch\n\nPretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful\nwhen you want to retrain a model on a custom chemistry domain rather than\nextending one of the released checkpoints. **Significantly more expensive than\n`kermt-continue-pretrain`** — no warm start, so the loss curves need to descend\nfrom scratch over many epochs.\n\n## Hardware requirements\n\nSame as `kermt-continue-pretrain`:\n\n- **GPUs**: 1–N CUDA-capable. The runner auto-detects via\n  `torch.cuda.device_count()`; `--gpus 0,2` overrides. Single-GPU fallback:\n  `--batch_size 32 --save_interval 500`. Multi-GPU keeps defaults\n  (`--batch_size 256` etc.). Note: `--gpus N` uses **torch.cuda** indexing,\n  which can differ from `nvidia-smi`'s display order on multi-GPU hosts\n  (PCI bus vs. CUDA enumeration). To target a specific physical GPU, set\n  `CUDA_VISIBLE_DEVICES` before invoking, or run\n  `python -c \"import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])\"`\n  to confirm which device you're picking.\n- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware\n  (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:\n\n  | GPU class                | VRAM       | Suggested `--batch-size` |\n  |--------------------------|------------|--------------------------|\n  | L4, T4, V100 16 GB       | 16–24 GB   | 32–64                    |\n  | A100 40 GB, L40, A40     | 40–48 GB   | 128                      |\n  | A100 80 GB, H100, H200   | 80 GB      | 256 (default)            |\n\n  These are rough starting points — pass `--batch-size N` to override.\n- **Disk**: tens of GB for shards + vocab + checkpoints, scaled by epochs.\n- **Wall time**: this is the big difference. Pretraining from scratch on an\n  11M-mol corpus at 100 epochs typically takes **days even on a multi-GPU box**.\n  The skill prints an estimate before launching; confirm with the user.\n\n## When to invoke\n\n- User wants to train a new model on a custom corpus (e.g. domain-specific\n  chemistry that the released ckpts don't cover).\n- User wants to reproduce a pretrain config end-to-end without depending on a\n  released ckpt.\n\nFor continuing an existing released ckpt, use `kermt-continue-pretrain`. For\nadding a cMIM decoder to an encoder-only grover_base ckpt, use\n`kermt-add-cmim-pretrain`.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — the pretrain corpus CSV with a `smiles` column. Single file\n  by convention; multi-file corpora deferred. Use `--val-csv` for a separate\n  validation set.\n- `--pretrain-target-mode {vocab|cmim|hybrid}` — which pretrain objective to\n  use. **No default** — must be set explicitly so the user makes an informed\n  choice:\n  - `vocab` — original GROVER-style atom + bond vocab prediction (encoder-only\n    output, lightweight).\n  - `cmim` — contrastive + SMILES reconstruction objective. Requires building\n    a SMILES vocab from the corpus.\n  - `hybrid` — both vocab and contrastive objectives jointly (the\n    state-of-the-art config from the KERMT manuscript).\n\nOptional:\n\n- `--val-csv <path>` — separate validation CSV. Without it, prepare_data\n  auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).\n- Training-hyperparameter overrides: `--epochs N` / `--batch-size N` /\n  `--init-lr F` / `--max-lr F` / `--final-lr F` / `--warmup-epochs F` /\n  `--weight-decay F` / `--dropout F` / `--save-interval N` / `--seed N`.\n  Anything not given is filled from `agent/config/defaults_pretrain.json`.\n- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /\n  `--contrastive-temperature F` (cmim and hybrid only).\n- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases\n  logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run\n  name is honored only alongside a project. Off by default.\n- `--gpus 0,2` — restrict to a GPU subset.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout.\n\n1. **Pre-flight: ensure container + system probe** (same as\n   `kermt-continue-pretrain` step 1). Refuse to proceed if `check_system`\n   reports gaps.\n\n2. **Compute run directory.**\n   ```\n   RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n   ```\n\n3. **Validate the corpus** (no ckpt to validate, so this is the only input\n   check):\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n       \"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>\"\n   ```\n   Abort on `ok: false`.\n\n4. **Prepare the data** — no vocab pass-through (we want fresh vocab from\n   corpus):\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/prepare_data.py --mode pretrain \\\\\n            --csv /data/<basename> --out /runs/data \\\\\n            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]\"\n   ```\n   Outputs land at `$RUN_DIR/data/prepare_data.json` with\n   `vocab_source: \"built_fresh\"`.\n\n5. **Estimate runtime + warn loudly.** This is critical for pretrain-from-scratch:\n   - \"Pretraining from scratch is days-scale even on multi-GPU; the released\n     KERMT checkpoints were each trained on millions of molecules for hundreds\n     of GPU-hours. If you mainly want to leverage existing knowledge for a\n     downstream task, consider `kermt-continue-pretrain` from a released ckpt\n     instead, which converges in hours instead of days.\"\n   - Show the corpus size × epochs × GPU count → estimated wall time.\n   - Ask for explicit confirmation unless `--yes` was given.\n\n6. **Launch the runner detached.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n       --name kermt-pretrain-scratch-<ts> \\\\\n       --run-dir $RUN_DIR -- \\\\\n       \"python agent/scripts/run_pretrain_local.py \\\\\n            --from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --out /runs \\\\\n            [--epochs N --batch-size N ...]\"\n   ```\n   Note: NO `--ckpt` flag (the runner refuses if both `--from-scratch` and\n   `--ckpt` are given). The runner uses the `arch` group from\n   `agent/config/defaults_pretrain.json` to size the model.\n\n7. **Report to the user.** Always include all of the following — do not\n   omit the TensorBoard line under output-length pressure:\n   - Container name + id\n   - `$RUN_DIR/run.json` (the manifest with `workflow: pretrain-scratch`,\n     `from_scratch: true`, `vocab_check: null`, `arch` from defaults, full\n     `cmd_replay`)\n   - Log file: `$RUN_DIR/logs/pretrain_ddp.log`\n   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir\n     $RUN_DIR/logs/tb`)\n   - Suggest `kermt-monitor <RUN_DIR>` for progress.\n\n## Hard rules\n\n- **Never accept a `--ckpt` flag.** From-scratch is exclusive with input\n  ckpt — the runner enforces this; the skill should too.\n- **Never silently default `--pretrain-target-mode`.** This is a significant\n  architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user\n  if not given on the CLI.\n- **Strong warning before launching.** From-scratch pretrain is the most\n  expensive workflow. The user needs to know what they're committing to.\n\n## Common errors\n\n- `--pretrain-target-mode is required when --from-scratch is set` → user\n  forgot the mode flag. Prompt.\n- `--from-scratch is incompatible with --ckpt` → user provided both; ask which\n  one they meant.\n- `defaults_pretrain.json has no arch group` → repo state issue (should never\n  happen on a fresh clone); points the user at running `kermt-setup` again.\n\n## What's in the manifest after a from-scratch run\n\nSame reproducibility fields as continue-pretrain (`repo.commit`, `kermt_image`,\n`cmd_replay`, `args_applied`), plus:\n\n- `workflow`: `\"pretrain-scratch\"`\n- `from_scratch`: `true`\n- `inputs.ckpt`: `null`\n- `ckpt_symlink`: `null`\n- `vocab_check`: `null` (not verified — vocab built from corpus is\n  authoritative for from-scratch)\n- `arch`: the values pulled from `agent/config/defaults_pretrain.json`'s\n  `arch` group (with any future CLI overrides applied).\n\n## Replayability\n\nSame as continue-pretrain: `cmd_replay` is a copy-pasteable command. If\n`ok_to_replay: false`, the kermt repo working tree was dirty at launch\ntime — check `repo.commit` and `git checkout` it first.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}