{"id":18076,"plugin_id":"plugins_6a84a979b5a081918922e037047514b5","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:36.456Z","digest":"9d5058bce318eb473f1e67ffe98cdb6ffe0e4d977e36de9f174fbe1cc18e3402","against":null,"payload":{"description":"Build a validated benchmax-sft-v1 dataset from existing labeled conversations, upload it, and explicitly launch a supervised finetuning run — no environment, rewards, or rollouts.","included_files":[],"name":"train-sft","skill_md_contents":"---\nname: train-sft\ndescription: Build a validated benchmax-sft-v1 dataset from existing labeled conversations, upload it, and explicitly launch a supervised finetuning run — no environment, rewards, or rollouts.\n---\n\n# Train with SFT\n\nUse this when the user **already has the completions they want the model to\nimitate** — support transcripts, labeled chat data, tool-call traces, input→output\npairs. There is no environment, tool loop, reward, or validate stage: the RL loop\nin this project's other skills does not apply. If the task needs the model to\n*discover* good behavior against a scorer, use the RL skills instead.\n\n## The whole flow\n\n```python\nfrom benchmax.sft import SftDataset, SftDatasetError\nfrom castform.platform import SftTrainingConfig, TrainerClient, upload_sft_assets\n\ntrain = SftDataset.from_jsonl(\"train.jsonl\")      # or SftDataset.from_rows(rows)\nuploaded = upload_sft_assets(dataset=train, run_name=\"support-sft\")\nrun_id = TrainerClient().launch_sft_run(\n    assets=uploaded,\n    name=\"support-sft\",\n    config=SftTrainingConfig(num_epochs=1, learning_rate=1e-5, seed=42),\n)\n```\n\nConstruction is **all-or-nothing**: `SftDataset` either satisfies the whole\n`benchmax-sft-v1` contract or raises `SftDatasetError` with every issue, ordered\nand line-aware. Fix the data the diagnostics point at — never pre-filter rows\nsilently or patch around individual issues without telling the user.\n\n## Row contract (one JSON object per line)\n\n```json\n{\n  \"messages\": [\n    {\"role\": \"user\", \"content\": \"What is 2 + 2?\"},\n    {\"role\": \"assistant\", \"content\": \"4\", \"weight\": 1}\n  ],\n  \"tools\": [],\n  \"metadata\": {\"id\": \"optional producer identity\"}\n}\n```\n\n- `system`/`user`: exactly `role` + non-empty string `content`.\n- `assistant`: optional string-or-null `content`, optional non-empty\n  `tool_calls`, optional integer `weight` `0 | 1` (omitted means `1`). Each\n  assistant turn needs content or a tool call; each row needs at least one\n  assistant turn with effective weight `1`.\n- `tool` results: exactly `role` + string `content` + non-empty `tool_call_id`;\n  every tool call gets exactly one result, in declaration order, before the\n  next non-tool message. Tool definitions are OpenAI function shapes;\n  `function.arguments` must decode as a JSON object.\n- Everything else is rejected: images/audio/multimodal parts, fractional\n  weights, legacy prompt/completion keys, unknown fields, duplicate JSON keys,\n  rows over 1 MiB, more than 1024 messages.\n\n**Masking:** set `weight: 0` on assistant turns that are context, not target\n(earlier drafts, retrieved answers, another model's output). Only weight-1\nturns contribute to the loss.\n\n## What the user actually chooses\n\n| arg | accepted | default |\n|---|---|---|\n| `num_epochs` | 1–100 | 1 |\n| `learning_rate` | (0, 0.1] | 1e-5 |\n| `max_context_tokens` | 256–8192, or 32768 / 65536 / 131072 (long context — must be enabled on the platform) | 8192 |\n| `save_interval` (steps) | 1–10000 | 20 |\n| `seed` | 0–2147483647 | 42 |\n| `lr_decay_style` | `\"constant\"` or `\"cosine\"` | unset — keeps the platform default |\n| `min_lr` | ≥ 0 and below `learning_rate` | unset |\n| `warmup_ratio` | 0–0.5 of total steps | unset |\n| `adam_beta2` | 0.9–0.999 | unset |\n| `grad_clip` | (0, 10] | unset |\n| `lora_rank` | 32 or 64 | unset — keeps the platform's fixed policy |\n| `global_batch_size` | 4–64, multiple of 4 at `max_context_tokens` ≤ 8192; 1–64 at a long-context rung (the platform picks the topology, so it decides divisibility) | unset — 4 |\n| `eval_interval` (steps) | 1–10000; only with an eval set | unset — derived from `save_interval` |\n\nEvery arg whose default reads `unset` is optional: leave it out and the\nplatform's own value applies, so an untouched config behaves exactly as before\nthese knobs existed. Do not set one just to restate a default.\n\n`lora_rank` picks the adapter's rank; alpha is always derived as 2x and is not\na knob. Rank 128 is not accepted — it trains, but serving cannot load it.\n\nModel (`Qwen/Qwen3.5-4B`) and GPU topology are platform-owned — do not invent\nknobs for them. A row that renders past `max_context_tokens` tokens fails the\nrun's preflight; trim long rows up front.\n\n## Held-out eval (optional)\n\nPass a second dataset to score during training:\n\n```python\nassets = upload_sft_assets(dataset=train, eval_dataset=held_out, run_name=\"...\")\n```\n\nThe eval set is scored on the live weights between training steps and plotted\nas `eval/loss` beside `train/loss`. Rules worth knowing before you offer it:\n\n- **At most 2048 rows**, under the same per-row limits as training rows. The\n  bound is a compute bound, not a taste one — eval runs on the training GPUs\n  and pauses training while it does.\n- **It is never billed.** Eval forward passes cost the user nothing, which is\n  also why the size and cadence are capped rather than left open.\n- **Cadence defaults to `save_interval`**, so by default every eval lands on a\n  checkpoint. Set `eval_interval` only to make it sparser; a much denser\n  cadence is rejected at launch.\n- An eval set is part of the data identity: changing it produces a different\n  upload prefix, and a resumed run must keep the one it started with.\n\nSkip it for small or exploratory runs — a held-out split costs training rows,\nand `train/loss` alone answers \"is this learning at all\".\n\n## Cost and consent\n\n`launch_sft_run` spends GPU credits. Ask the human before calling it, every\ntime — preparing and uploading the dataset first is free and fine. Steps per\nepoch ≈ rows / 4 (tiny datasets pad by repeating their first rows, so 1–3-row\ndatasets train on repeats; prefer at least a few dozen rows). Stopping a run\nkeeps only checkpoints already uploaded; work since the last one is lost.\n\nIf launch fails with \"SFT launch is not enabled\", the platform gate is off for\nthis account — surface that to the user rather than retrying.\n\n## Monitor\n\nSame as any run (`view-progress` skill): `castform runs status <id>`, and\n`castform runs scalars <id>` — watch `train/loss` fall, and `eval/loss` too\nwhen the run has an eval set. There are no rollouts or reward curves for SFT\nruns. The run page shows loss, dataset prefix, and config; runs with an eval\nset also get an eval tab, chart-only.\n\n## Reference\n\nThe canonical worked example (streaming a pinned public corpus, bounded\nmapping, offline tests, explicit paid `--launch` gate) ships in the benchmax\nrepo under `examples/sft/pii_masking/`.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}