{"id":17591,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.533Z","digest":"612f361d20a806b5eb9b847f788ac06ccf78910fa1f954d203c56f582cde10ae","against":null,"payload":{"name":"kermt-finetune","description":"Finetune a pretrained KERMT encoder on a labeled CSV. The skill validates the input checkpoint (must be a pretrain ckpt — grover_base / cmim / hybrid), validates the labeled CSV, prepares the data (clean + features + optional split), then launches main.py finetune inside the kermt container (detached for hours-scale runs). Hyperparameters come from agent/config/defaults_finetune.json with per-flag CLI override.","included_files":[],"skill_md_contents":"---\nname: kermt-finetune\ndescription: Finetune a pretrained KERMT encoder on a labeled CSV. The skill validates the input checkpoint (must be a pretrain ckpt — grover_base / cmim / hybrid), validates the labeled CSV, prepares the data (clean + features + optional split), then launches main.py finetune inside the kermt container (detached for hours-scale runs). Hyperparameters come from agent/config/defaults_finetune.json with per-flag CLI override.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: workflow-skill\n  risk_tier: skill\n# Line/token budget: targets ~250 lines / ~3000 tokens — within the\n# 500-line / 5000-token cap for skill files.\n---\n\n# kermt-finetune\n\nFinetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill\nis the workflow orchestrator: validate ckpt, validate data, prepare data,\nlaunch the runner detached, return a run directory + container name.\n\n## Hardware requirements\n\n- **GPUs**: 1 by default (single-GPU); pass `--gpus 0` (or whichever id) to\n  select one. For faster training on a multi-GPU host, pass `--num-gpus N`\n  (N>1) to run data-parallel DDP across N GPUs — `--batch-size` is then\n  per-GPU (effective global batch = batch_size × N).\n- **VRAM**: ≥ 8 GB for the default `batch_size 32` configuration. Lower VRAM\n  works at smaller batch sizes — pass `--batch-size N` to override.\n- **Disk**: a few GB per run (checkpoint + features + logs).\n- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).\n  `kermt-setup` validates this up-front.\n\n## Inputs\n\nRequired:\n\n- `--csv <path>` — labeled CSV. First column is `smiles`; every other column\n  is a target.\n\nCheckpoint (optional — defaults to the released model if omitted):\n\n- `--ckpt <path>` — input pretrain checkpoint (grover_base / cmim / hybrid).\n  The validator refuses already-finetuned ckpts with a redirect to\n  `kermt-infer`. **If omitted**, the skill offers to download the released\n  pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and finetune from it —\n  see \"Resolve & validate the checkpoint\" (workflow step 3).\n- `--pretrained-release` — explicit opt-in to use the released model without\n  the interactive prompt (for non-interactive / agent runs). Mutually\n  exclusive with `--ckpt`.\n- `--model-dir <dir>` — where to save the downloaded bundle (default\n  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is\n  reused, not re-downloaded.\n\nOptional:\n\n- `--dataset-type {regression | classification | multiclass}` — default\n  `regression` (from `defaults_finetune.json`). Drives loss, metric defaults,\n  and head initialization. For classification tasks pass\n  `--dataset-type classification`.\n\n- `--targets COL [COL ...]` — explicit target column names. If omitted, the\n  validator auto-detects numeric non-smiles columns and the skill confirms\n  with the user before proceeding.\n- `--val-csv <path>` and `--test-csv <path>` — user-provided val + test\n  splits. Either pass both or pass neither (the skill auto-splits using the\n  configured `--split-type`).\n- `--split-type {random | scaffold_balanced | index_predetermined}` —\n  default `scaffold_balanced` from `defaults_finetune.json`.\n  - `random` and `scaffold_balanced`: build the val/test split internally\n    from the train CSV. No `--val-csv` / `--test-csv` needed.\n  - `index_predetermined`: **requires** pre-split CSVs passed via\n    `--val-csv` + `--test-csv` (and, separately, per-fold index files —\n    see `kermt/util/utils.split_data`). Use this when the dataset ships\n    its own canonical split (e.g. `tests/data/Biogen_for_grover/scaffold/\n    balance/<endpoint>/{train,val,test}.csv`).\n- `--metric NAME` — `mae` (regression default), `auc` (classification default),\n  or any name `kermt.util.metrics.get_metric_func` accepts.\n- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /\n  `--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /\n  `--bond-drop-rate F` / `--dist-coff F` / `--early-stop-epoch N` /\n  `--seed N` — training-hyperparameter overrides. Anything not given is\n  filled from `agent/config/defaults_finetune.json`.\n- `--ffn-hidden-size N` / `--ffn-num-layers N` — shared FFN trunk dims.\n- `--ffn-num-task-specific-layers N` / `--ffn-task-specific-hidden-size H` —\n  per-target FFN heads (default 0 = off; useful for heterogeneous multi-target\n  finetunes). Both must be set together when N > 0.\n- `--ensemble-size N` / `--num-folds N` — multi-model / k-fold CV. Default 1\n  each.\n- `--gpus 0` — single GPU id for single-process finetune (default 0). Ignored\n  when `--num-gpus > 1`.\n- `--num-gpus N` — number of GPUs for data-parallel DDP finetune. Default 1\n  (single-process, unchanged). N>1 runs `main.py finetune` with `WORLD_SIZE=N`\n  (one process per GPU); `--batch-size` is per-GPU.\n- `--from-prepare <dir>` — skip the prepare step and reuse an existing\n  `prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.\n\n## Workflow\n\nLet `$KERMT_REPO` be the path to your kermt repo checkout, and assume\n`kermt-setup` has built `kermt:latest`. All paths below are on the host; the\nhelper bind-mounts them at known container paths.\n\n1. **Pre-flight: ensure container + system probe.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c \"\n   import json, sys; d = json.load(sys.stdin)\n   if not d['ok']:\n       print('System check failed:', d['gaps']); sys.exit(1)\n   print(f'OK: {len(d[\\\"gpus\\\"])} GPU(s); CUDA via container toolkit')\n   \"\n   ```\n   Refuse to proceed if `ok: false`.\n\n2. **Compute run directory.**\n   ```\n   RUN_DIR=$KERMT_REPO/runs/finetune_$(date -u +%Y-%m-%dT%H-%M-%SZ)\n   ```\n\n3. **Resolve & validate the checkpoint.**\n\n   **Resolve — only if `--ckpt` was omitted.** Default to the released\n   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:\n   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:\n     \"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2\n     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)\n     and finetune from it? [y/N]\". **Never download without an explicit yes**\n     (or `--pretrained-release`). If both `--ckpt` and `--pretrained-release`\n     are given, abort — they conflict.\n   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor\n     `--model-dir <dir>` if given. An already-complete bundle is reused.\n   - **Download** (foreground; ~282 MB on first fetch):\n     ```\n     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \\\n         \"python agent/scripts/fetch_released_model.py --out /model\"\n     ```\n     Parse the JSON; abort on `ok: false` (surface `errors`). On success set\n     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.\n\n   **Validate** the resolved (or user-provided) ckpt:\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \\\n       \"python agent/scripts/check_checkpoint.py --mode finetune_init --ckpt /ckpt\"\n   ```\n   Parse the JSON. Abort on `ok: false`. The validator rejects already-\n   finetuned ckpts (`has_task_ffn: true`) with a redirect to `kermt-infer`.\n\n4. **Validate the data.**\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \\\n       \"python agent/scripts/check_data.py --mode finetune --csv /data/<basename> [--targets COL1 COL2 ...]\"\n   ```\n   If `--targets` was not given by the user, surface `auto_detected_targets`\n   from the JSON and ask the user to confirm before continuing. Abort on\n   `ok: false`.\n\n5. **Prepare the data** (skip if `--from-prepare` given).\n\n   **Pre-flight: check for sibling val.csv / test.csv.** Before invoking\n   prepare_data, inspect the parent directory of `<user-csv>`. If a\n   canonical-looking sibling `val.csv` (or `val_*.csv` — common variants\n   include `val_T.csv`, `val_clean.csv`) AND a matching `test.csv` /\n   `test_*.csv` exist next to the train CSV, the dataset ships its own\n   pre-defined split. **In that case set `--split-type index_predetermined`\n   AND pass `--val-csv` / `--test-csv`** — otherwise the configured\n   `split_type` (default `scaffold_balanced`) will re-split the train CSV\n   from scratch and silently discard the user's val/test files. When in\n   doubt — or when the sibling files use non-canonical suffixes (`_T`,\n   `_v2`, etc.) — surface the situation to the user and ask which they\n   want.\n\n   **Quoting target names.** If any of the `--targets` column names\n   contain shell metacharacters (`>`, `&`, `|`, `(`, `)`, `$`, etc.),\n   single-quote each one when passing on the CLI to keep the shell from\n   eating part of the name. Example: `--targets 'Log_Caco2_Papp_A>B'\n   'logD'`. The CSV header itself is read directly by the downstream\n   trainer and is unaffected, but the prepare_data.json manifest's\n   `targets[]` field captures whatever the shell delivers — unquoted\n   metacharacters get truncated there.\n\n   **Mount note:** `kermt_container.sh --data <host-csv>` mounts the\n   parent directory of `<host-csv>` at `/data`. `--val-csv` and\n   `--test-csv` must therefore reference files in that same parent\n   directory. If val/test live in a separate directory (e.g. a sibling\n   `splits/` folder), mount the parent of all three using `--data <dir>`\n   on a directory rather than a file.\n\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \\\n       \"python agent/scripts/prepare_data.py --mode finetune \\\\\n            --csv /data/<basename> --out /runs/data \\\\\n            --split-type <split_type> \\\\\n            [--val-csv /data/<val-basename> --test-csv /data/<test-basename>] \\\\\n            [--val-frac 0.1 --test-frac 0.1 --seed 0] \\\\\n            --targets <COL1> [COL2 ...]\"\n   ```\n   Outputs land at `$RUN_DIR/data/prepare_data.json`. For `scaffold_balanced`\n   and `index_predetermined`, prep emits a single `clean_full_csv` + `.npz`;\n   the runner passes them through to `main.py finetune` which calls\n   `split_data` internally with the user-supplied seed.\n\n6. **Estimate runtime + echo applied defaults.**\n   - Finetune wall time is typically minutes-to-hours on 1 GPU.\n   - Surface a summary of every flag that was filled from the defaults\n     vs user-supplied, so the user knows what was assumed. The runner\n     records this in `args_applied`.\n   - Sample message:\n     `\"Filling from defaults_finetune.json: epochs=30, batch_size=32,\n       split_type=scaffold_balanced. Override any of these with --<flag>.\"`\n\n7. **Targets confirmation gate (hard requirement).** Before launching the\n   runner, regardless of how the targets list was determined (CLI `--targets`,\n   auto-detection in step 4, or a user natural-language request like\n   \"finetune on Caco2 and HLM\"), echo the final targets list to the user with\n   an explicit count:\n   `\"Will finetune on N target(s): COL1, COL2, ...\"`. If the user's request\n   specified a subset that doesn't match this list (e.g., they asked for 2\n   tasks via natural language but the list still has 4), treat it as a\n   discrepancy and re-prompt with the diff — never silently proceed on the\n   wrong target set. Wait for explicit confirmation before launching unless\n   `--yes` was given.\n\n8. **Launch the runner detached.** (Consistent with the pretrain skills.)\n   ```\n   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\\\\n       --name kermt-finetune-<ts> \\\\\n       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\\\\n       \"python agent/scripts/run_finetune_local.py \\\\\n            --ckpt /ckpt \\\\\n            --prepare-manifest /runs/data/prepare_data.json \\\\\n            --dataset-type <type> \\\\\n            --out /runs \\\\\n            [--gpus 0] \\\\\n            [--num-gpus N] \\\\\n            [--epochs N --batch-size N --init-lr F ...] \\\\\n            [--ffn-num-task-specific-layers N --ffn-task-specific-hidden-size H]\"\n   ```\n   Returns the container name + id + log file path.\n\n9. **Report to the user.** Output a short summary:\n   - Container name + id\n   - `$RUN_DIR/run.json` (manifest with cmd_replay + image digest)\n   - Log file: `$RUN_DIR/logs/finetune.log`\n   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir\n     $RUN_DIR/logs/tb`)\n   - Final checkpoints land at `$RUN_DIR/ckpt/fold_0/model_0/model.pt`\n     (best-val) and `last_checkpoint.pt` (sibling, auto-resume target).\n     Held-out test predictions + metrics land at\n     `$RUN_DIR/ckpt/fold_0/test_result.csv`. Paths vary with `--num-folds`\n     / `--ensemble-size`.\n   - To follow progress: `kermt-monitor <RUN_DIR>` (one-shot) or\n     `docker logs -f <container-name>` (streaming).\n   - To block until the run finishes (useful for short test runs):\n     `docker wait <container-name>` — prints the exit code on completion.\n\n## Hard rules\n\n- **Never download the released model without consent.** When `--ckpt` is\n  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user \"yes\"\n  or an explicit `--pretrained-release` flag. `--ckpt` and\n  `--pretrained-release` are mutually exclusive.\n- **Never modify the user's input ckpt.** The runner passes its path via\n  `--checkpoint_path`; `task/train.py` loads it read-only into the model and\n  attaches a new FFN head. The source file stays untouched.\n- **Arch comes from the ckpt, not from CLI/defaults.** The runner extracts\n  `hidden_size`, `depth`, `num_attn_head`, `activation`, `embedding_output_type`,\n  `self_attention` (+ `attn_hidden` / `attn_out` when applicable) from the\n  ckpt's saved_args. There is no `--hidden-size` flag on this runner.\n- **Never block on the long-running finetune.** The skill launches via\n  `run_detached` and returns immediately after step 9. Use `kermt-monitor`.\n- **Echo applied defaults back to the user.** The `args_applied` field of\n  `run.json` records every flag's value + source (user / default-config).\n  Surface a one-line summary of every filled-from-default flag so the user\n  knows what was assumed.\n\n## Common errors\n\n- `finetune_init requires a pretrain ckpt (grover_base / cmim / hybrid)` →\n  the ckpt you passed is already finetuned (has task FFN heads). Pick a\n  pretrain ckpt instead, or use `kermt-infer` if you want to run\n  predictions with the existing finetuned model. To resume a finetune on\n  the SAME dataset, bypass the skill and call\n  `python main.py finetune --checkpoint_path <ckpt> ...` directly — the\n  agent skill doesn't support resume because saved-task identity\n  can't be machine-verified against the new training data.\n- `prepare_data manifest reports ok=False` → check `errors` for the failed\n  step (typically clean_smiles or save_features). Fix and re-run.\n- `ffn_num_task_specific_layers=N>0 but ffn_task_specific_hidden_size is unset`\n  → MTL heads need an explicit hidden size. Pass `--ffn-task-specific-hidden-size H`.\n- `finetune is single-GPU` (from `--gpus 0,1`) → `--gpus` selects one device\n  for single-process finetune. For multi-GPU, use `--num-gpus N` (DDP) instead.\n\n## Replayability\n\nThe `run.json` `cmd_replay` field is a single-line command that re-runs the\nfinetune with the same inputs, hyperparameters, and arch. To replay inside\nthe kermt container:\n\n```bash\n$(jq -r .cmd_replay $RUN_DIR/run.json)\n```\n\nIf `ok_to_replay: false` in the manifest (because the kermt repo working\ntree was dirty at launch time), the replay may not be bit-exact — pin the\nexact commit via the `repo.commit` field and `git checkout` it\nfirst.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}