← ApprenticeCONTENT HISTORY

Update to Apprentice

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "apprentice-deploy",
  "description": "Use ONLY when a user explicitly asks to serve or deploy a model already fine-tuned with Apprentice: \"serve this model\", \"deploy the adapter\", \"run it in my cluster\", \"vLLM\", \"Kubernetes manifests for it\". Writes vLLM Deployment and Service manifests into the user's repo mirroring existing conventions, and covers serving MLX adapters on a Mac. Never volunteer deployment while capturing calls, optimizing a prompt, or training: delegate those to apprentice-capture and apprentice-train.",
  "included_files": [
    {
      "relative_path": "references/deploy-kubernetes.md",
      "size_in_bytes": 2097
    }
  ],
  "skill_md_contents": "---\nname: apprentice-deploy\ndescription: >\n  Use ONLY when a user explicitly asks to serve or deploy a model already\n  fine-tuned with Apprentice: \"serve this model\", \"deploy the adapter\", \"run\n  it in my cluster\", \"vLLM\", \"Kubernetes manifests for it\". Writes vLLM\n  Deployment and Service manifests into the user's repo mirroring existing\n  conventions, and covers serving MLX adapters on a Mac. Never volunteer\n  deployment while capturing calls, optimizing a prompt, or training:\n  delegate those to apprentice-capture and apprentice-train.\nlicense: MIT\n---\n\n# Serve a fine-tuned model\n\nOnly when asked. A user uploading rows or running an optimize job has not asked to deploy\nanything, and offering manifests there is noise at best. This skill exists as its own trigger\nso deployment instructions stay out of sessions that are about data.\n\n**Asking about deployment is not asking for files.** \"How would I serve this?\", \"what does\nvLLM need?\", \"is my cluster big enough?\" are questions: answer them, write nothing. Write\nmanifests only when the user asks for the files, in words like \"write the manifests\", \"add\nthe Deployment\", \"set this up in my repo\". When it is genuinely ambiguous, say what you are\nabout to create and where, and let the user say go. The plugin declares Write, so this gate\nis the only thing standing between a question and a commit.\n\nThe precondition, surfaced before anything else: the fine-tune has passed its eval, and for\nthe cluster path there is **at least one GPU node**.\n\n## In the user's own Kubernetes cluster\n\nInference stays inside the user's network. Read `references/deploy-kubernetes.md` and the published\n[Kubernetes guide](https://docs.runapprentice.com/how-to/deploy-kubernetes), then write the vLLM Deployment and Service into the\nrepo, mirroring the conventions already there (the user's registry, namespace, ingress\npattern, label scheme), and state the honest GPU sizing.\n\nNever invent cluster names, namespaces, or registries. With no existing manifests visible, ask\nfor one rather than guessing.\n\n## On a Mac\n\nMac-trained MLX adapters are served on the Mac with `mlx_lm.server`\n([docs](https://docs.runapprentice.com/how-to/deploy-mlx)).\n\nDo not claim an MLX adapter can be served by vLLM elsewhere. That conversion path has no\npublished, verified recipe, and saying otherwise sends a user down a road that dead-ends.\n\n## What deployment does not do\n\nServing a model is not promotion. Activating a model in Apprentice records the promotion in\nthe console and leaves serving unchanged; routing live production traffic to the smaller model\nis still in development. A user who deploys still decides what calls the endpoint.\n"
}

SHA-256: f9ca7b148b86daa8aa61d8d8599e080b0c6e763efdcda7ba9f5fd4a9f64b1bd5