← ApprenticeCONTENT HISTORY

Update to Apprentice

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Use when a user is ready to fine-tune a small model and already has verified rows, or asks what training needs: says \"fine-tune\", \"distill\", \"train a small model\", \"LoRA\", or asks how many rows training takes. Also use for drift on a model already running: whether it still holds quality and when to retrain. Local MLX training on Apple silicon is the path that works; hosted training is not shipped yet. For a user still deciding whether a cheaper model could work, or without a dataset, use apprentice instead, and delegate recording calls and prompt optimization to apprentice-capture and serving to apprentice-deploy.",
  "included_files": [],
  "name": "apprentice-train",
  "skill_md_contents": "---\nname: apprentice-train\ndescription: >\n  Use when a user is ready to fine-tune a small model and already has\n  verified rows, or asks what training needs: says \"fine-tune\", \"distill\",\n  \"train a small model\", \"LoRA\", or asks how many rows training takes. Also\n  use for drift on a model already running: whether it still holds quality\n  and when to retrain. Local MLX training on Apple silicon is the path that\n  works; hosted training is not shipped yet. For a user still deciding\n  whether a cheaper model could work, or without a dataset, use apprentice\n  instead, and delegate recording calls and prompt optimization to\n  apprentice-capture and serving to apprentice-deploy.\nlicense: MIT\n---\n\n# Fine-tune a small model\n\nTraining is the second half of the product. The first half, a dataset of verified rows, has\nto exist before any of this runs.\n\n## Gold only. This is the part users hit first\n\nTraining reads **gold rows only**: rows a human verified. Silver counts for prompt\noptimization and never for training.\n\n| | Rows needed |\n|---|---|\n| Hosted optimize | 20 verified (gold + silver) |\n| Local training | enough gold to fill a batch after the held-out split |\n| Hosted training | 500 gold, and not yet available (see below) |\n\nA user capturing through an API key accumulates raw traces, and a user uploading rows\naccumulates silver. Neither is gold, because recording a verdict needs a signed-in console\nsession. So the honest first answer to \"can I train?\" is usually\n\"not yet, and here is the gap\": count the gold rows, name the number missing, and link\n`https://runapprentice.com/tasks/<task>/review`.\n\nDo not auto-approve rows to reach 500. Gold means a human checked it. A model grading its own\noutput is not verification, and every eval downstream inherits the lie.\n\n## Local, on Apple silicon. This is the path that works today\n\nHosted training is **still being built**, which the\n[docs index](https://docs.runapprentice.com/llms.txt) states plainly: the shipped features are\nprompt optimization and local training. `client.train(task)` accepts a job, and the server\nraises `NotImplementedError` because no GPU trainer is configured. Do not present it as\nworking, and do not send a user to it.\n\nLocal training is real, free, and runs on the user's own Mac via MLX. No Apprentice account is\nneeded when a CSV is supplied:\n\n```\napprentice train <task> --local --data examples.csv --effort high\n```\n\nOr from Python:\n\n```python\nclient.train_local(\"duplicate-search\", data=\"examples.csv\", effort=\"high\")\n```\n\nThe model is `Qwen3.5-4B` (Apache-2.0), LoRA fine-tuned.\n\n**`--effort` is a memory profile, not a quality knob.** Every tier trains the same model on\nthe same examples for the same number of passes; a higher tier uses a bigger batch to finish\nsooner and needs more memory. It is auto-detected from the Mac's unified memory. Long prompts\nneed more memory per example, so drop a tier on an out-of-memory error.\n\nSay this plainly when a user asks whether `--effort low` gives a worse model. It does not.\n\nRun history: `train --local` reaches the console when an API key is set, and never otherwise.\n\n## Reading the result\n\nThe number that matters is the held-out score, not the training loss. Held-out rows are the\nones the model never saw, and the eval gate uses gold only.\n\nTwo public runs for scale, reproducible in\n[apprentice-benchmark](https://github.com/singhabhishekkk/apprentice-benchmark):\n\n- Receipt extraction (OCR text from 200 real scanned receipts, field-level F1, seed 42): a\n  LoRA fine-tuned Qwen3.5-4B reaches 89.2 on a 60-row held-out split, against GPT-4o-mini at\n  72.9 plain and 84.2 after GEPA.\n- JSON extraction (100 rows, same metric): the fine-tuned 4B reaches 88.9 on a 30-row\n  held-out split, against 83.1 plain and 85.6 after GEPA.\n\nThose held-out splits are small, so treat them as directional. The user's own run on the\nuser's own data is the number that decides anything.\n\nNever promise the small model wins. It might not, and the eval exists to find that out.\n\n## Promotion, stated accurately\n\nActivating a model **records** the promotion in the console. It does not move production\ntraffic: serving is unchanged, and routing live traffic to the smaller model is still in\ndevelopment. Say this rather than implying a takeover the product does not perform.\n\n## Drift, once a model is live\n\nCapture records the call, feedback records whether it worked, and that feedback score is what\nthe console's Drift view charts. Without feedback a user sees traffic volume and no quality\nsignal, which is the common reason \"is it still good?\" has no answer.\n\nThe retrain trigger is a real drop in that score on real traffic, not a calendar date.\n\n## Serving\n\nOnce a fine-tune passes its eval, use `apprentice-deploy`.\n"
}

SHA-256 of public snapshot: 0d1d6b6d2d09cf6ca550c3ca017cfbc41224ab4480d1fdc5a02857e7a8acea81