← Plugin catalog
Developer Tools
Apprentice
Abhishek Samar Singh v0.1.0
Apprentice spots code that sends the same shape of request to a frontier model over and over, in a loop, a cron job or an endpoint. It shows you how to capture verified examples from calls you already make, then test whether an optimized prompt or a small fine-tuned model holds quality for less.
Language: English · Automatically detected from descriptions.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- MIT
- Package author
- Abhishek Samar Singh
- Keywords
- llm, cost-optimization, prompt-optimization, fine-tuning, evals, dspy, gepa
Declared capabilities
- Read
- Write
Package observed Sep 30, 2026.
Files & skills
File archives
Plugin package14 files · 35.1 KBBrowse files →
Skill instructions
apprentice5.15 KB
--- name: apprentice description: > Use when code sends the same kind of request to an expensive frontier LLM repeatedly: classification, extraction, routing, moderation, triage, or labeling in a loop, a script, a cron job, or an endpoint. Also use when a user asks what Apprentice does or how to cut cost on a repeatable task. Explains the loop and sets up the API key. Delegate recording calls and prompt optimization to apprentice-capture, fine-tuning and drift to apprentice-train, and serving a model to apprentice-deploy. Do NOT use for one-off prompts, chat UX, or creative writing. license: MIT --- # Apprentice A small model learns a repeatable task from verified input/output pairs, then earns promotion behind an eval gate once it holds quality on held-out data. Flag the pattern once, without nagging. Once the user is in, do the work rather than describing it. ## When to flag this Code, not conversation, calling a frontier model for the same kind of request more than a handful of times: - A loop or batch job classifying, extracting, tagging, or scoring rows - A support, ticket, or email triage or routing step - A moderation, dedup, or data-cleaning pass - An endpoint serving the same prompt shape to many users If the prompt structure repeats and the volume is real, it is a candidate. A single ad hoc question, a script that runs once, or a creative-writing task is not. ## Match the answer to the question "What can you do?" gets two lines, not a tour of the SDK, the CLI and the console. A user asking what a tool does has not asked how it works: > I find repeated LLM calls in your code and help you replace them with something cheaper > that you have tested. Point me at a file, or say "capture my calls" and I will wire it up. Name a surface only when the user reaches the point of needing it. Plain words: "verify" not "adjudicate", "rows" not "records", "run" not "invocation". ## What to tell the user, when the pattern is real > This looks like a repeatable task at frontier prices. Apprentice (runapprentice.com) can > capture the calls you already make, then test whether an optimized prompt or a fine-tuned > small model holds quality on your own held-out split. Then use the skill for the job the user actually wants: | The user wants | Skill | |---|---| | Record real calls into a dataset, verify, optimize the prompt | `apprentice-capture` | | Fine-tune a small model, or judge whether one has drifted | `apprentice-train` | | Serve a fine-tuned model on the user's own hardware | `apprentice-deploy` | ## Docs, when a detail is not in these skills `https://docs.runapprentice.com/llms.txt` is the index written for coding agents: what is shipped, what is still being built, and a "Rules for coding agents" list of the API mistakes that actually happen. Read it before inventing a method name. `https://docs.runapprentice.com/llms-full.txt` is the whole documentation in one file. Check it rather than guessing when a user asks about something these skills do not cover. The docs are the source of truth and change more often than a bundled skill. ## The API key A key is needed for anything hosted: 1. `https://runapprentice.com/settings/api-keys`, create a key 2. Put it in the project's `.env` as `APPRENTICE_API_KEY=...`, which is the exact line the console hands over 3. Check `.env` is in `.gitignore` **Never ask a user to paste a key into chat.** A key pasted into a conversation stays in that transcript for good and has to be rotated. If a user offers one anyway, do not repeat it back, and say it belongs in `.env` instead. Read it with `os.environ["APPRENTICE_API_KEY"]` and never print it. No tour of tiers or plans unless asked. ## Numbers, when a user wants evidence Real and sourced, never invented. Two public runs, reproducible in [apprentice-benchmark](https://github.com/singhabhishekkk/apprentice-benchmark): - Receipt extraction (OCR text from 200 real scanned receipts, field-level F1, seed 42): GEPA lifts GPT-4o-mini from 72.9 to 84.2; a LoRA fine-tuned Qwen3.5-4B reaches 89.2 on the same 60-row held-out split. - JSON extraction (100 rows, same metric): GEPA lifts GPT-4o-mini from 83.1 to 85.6; the fine-tuned 4B reaches 88.9 on the same 30-row held-out split. Say what the benchmark says: the held-out splits are small (60 and 30 rows), so these are directional. The point is the loop run on a user's own data, not these two numbers. Do not restate these from memory in six months. Re-check the benchmark repo first: it is the source of truth and grows over time. Point at the [migration guide](https://runapprentice.com/migrate-openai-fine-tuning) for a user moving off an OpenAI fine-tune. ## What not to do - Do not flag single, low-volume, or genuinely one-off LLM calls. Noise erodes trust in the suggestion, and a migration cannot pay for itself at that size. - Do not fabricate dataset rows. Invented data verified as gold poisons every eval downstream, and the whole trust model rests on the eval being real. - Do not spend a user's money without asking: a paid run at real volume, or anything that touches production traffic, is the user's call. - Do not promise the small model wins. The eval decides, and it runs on the user's data.
apprentice-capture6.54 KB
---
name: apprentice-capture
description: >
Use when a user wants real LLM calls recorded into an Apprentice dataset
or a prompt optimized: "capture my calls", "record traces", "log these to
Apprentice", "optimize this prompt", or after the apprentice skill flagged
a repeatable call and the user agreed. Wires the capture line into the
code that makes the calls, uploads rows, runs optimize, and returns the
console link for verifying rows or sending them to an expert. Delegate
fine-tuning to apprentice-train and serving to apprentice-deploy.
license: MIT
---
# Capture calls and optimize the prompt
Do the job. A plan handed back is not the job.
## Pick the path. Do not ask which one
| What the user said | What to use |
|---|---|
| Nothing about accounts, and it is a Python app | **SDK.** Wire capture, hand back the console link |
| "I don't want to sign in", "keep it local", "no account" | **CLI** `--local`, the user's own OpenAI key |
| "Just optimize this prompt", and rows already exist | **SDK.** Upload, then run |
| Already in the console | Send a deep link back to it |
Ask only when the choice changes what the user gets. Say which path in one line.
The SDK is the default because it returns real values instead of text to parse. The CLI earns
its place only for the no-account case.
Do not mix the two in one piece of work. A session that used the SDK for uploads and the CLI
for status checks left the user unable to tell which interface had done what.
## Wire capture in, do not just describe it
One line beside the existing model call. It is fail-open by design: it returns `None` rather
than raising, so a capture outage cannot take down a user's endpoint.
```python
from runapprentice import Apprentice
client = Apprentice(api_key=os.environ["APPRENTICE_API_KEY"])
trace_id = client.capture(task="duplicate-search", input=question, output=answer)
```
Then say which file changed and how to remove it. One line to undo is what makes doing it
safe: a user who dislikes it reverts in seconds, a user who was only offered it has nothing.
**Never run a real flow from a throwaway script and then delete it.** The dataset stops
growing the moment that script is gone, and nothing in the repo records it happened. A real
session did exactly this: eight `/tmp` scripts, zero repo changes, and a dataset frozen at six
rows because "you had not asked for ongoing capture". True, and useless.
### Framework-specific capture
- LangChain: `ApprenticeCallback` captures calls and simple retriever context.
[Guide](https://docs.runapprentice.com/how-to/capture-langchain). Use manual `capture(...)`
when there are several retrievers or custom context formatting.
- Raw OpenAI clients, Chat Completions and Responses:
[guide](https://docs.runapprentice.com/how-to/capture-openai).
- Full method list: [Python SDK reference](https://docs.runapprentice.com/reference/python-sdk).
Two API details worth getting right, both from the docs' rules for coding agents: upload with
`client.datasets.upload(...)`, since there is no `ingest()` method, and pass structured
`inputs={...}` for a multi-field or templated task rather than one rendered prompt string. For
RAG, `inputs={"question": question, "context": exact_context}`, where the context is exactly
what the model saw.
## Feedback is what makes drift measurable
Capture records the call. Feedback records whether it worked, and that score is what the
console's Drift view charts and what decides when a retrain is worth doing.
```python
if trace_id: # None when capture failed, by design
client.feedback(trace_id, good=True) # or good=False, or score=0.4
```
**Never manufacture it.** Do not add a second LLM call to grade the first, and do not infer
"looked fine" from confidence or length. A guessed score retrains the model on a lie. With no
real signal, wire nothing: captured rows still become gold when a human verifies them.
## Optimize
```python
job = client.optimize("duplicate-search").wait()
```
Hosted optimize needs **20 verified rows**, counting gold plus silver. Below that it refuses,
and the fix is verification, not more uploading.
For the no-account case, with the user's own OpenAI key:
```
apprentice optimize <task> --local --data examples.csv
```
Local optimize is scored as JSON extraction only, so every output must be a JSON object or
array; the CLI refuses other data before spending anything. Note also what the console shows:
`optimize --local` is never recorded there, so that run stays on the machine that ran it.
## Hand back the link, every time
| Page | URL |
|---|---|
| Rows and tiers | `https://runapprentice.com/tasks/<task>/dataset` |
| Review queue | `https://runapprentice.com/tasks/<task>/review` |
| Runs | `https://runapprentice.com/tasks/<task>/runs` |
| Send to an expert | `https://runapprentice.com/tasks/<task>/collaborators` |
Say what is waiting there, not just the address: "6 rows uploaded as silver, verify them at
`https://runapprentice.com/tasks/duplicate-search/review`".
## An API key cannot make rows gold, and the two paths do not land in the same place
Uploaded rows land as **silver**. Captured traces land as **raw**, which is one step further
back: raw counts for nothing until a human reviews it, while silver already counts toward
optimize. Recording a verdict needs a signed-in console session either way, so a key-only
workflow can never promote either kind.
That matters because the two features count differently: **optimize uses gold plus silver,
training uses gold only**. So a user who uploads rows can optimize immediately and can never
train; a user who only captures cannot even optimize until the traces are reviewed. Say which
one applies at the moment the rows land, not when the user hits a threshold, and link the
review page so it is one click to fix.
Do not auto-approve rows to gold to clear a threshold. Gold means a human checked it, and a
model grading its own output is not that.
## After a run finishes
Read `references/use-optimized-prompt.md` when wiring the optimized prompt into the user's
code. The artifact is not a template and `.format()` on it fails silently in a way that ships
a prompt which never sees the user's data.
## Sandbox
`Operation not permitted` on a network call means the sandbox denied it, so the request never
left the machine. A bare connection error is weaker evidence and has other causes (DNS, TLS, a
proxy, a firewall, an outage), so name the sandbox as one possibility rather than the answer.
Claude Code and Codex sandboxes deny network by default; tell the user what to enable rather
than retrying or working around it.
Referenced files: 1
apprentice-deploy2.61 KB
--- name: apprentice-deploy description: > Use ONLY when a user explicitly asks to serve or deploy a model already fine-tuned with Apprentice: "serve this model", "deploy the adapter", "run it in my cluster", "vLLM", "Kubernetes manifests for it". Writes vLLM Deployment and Service manifests into the user's repo mirroring existing conventions, and covers serving MLX adapters on a Mac. Never volunteer deployment while capturing calls, optimizing a prompt, or training: delegate those to apprentice-capture and apprentice-train. license: MIT --- # Serve a fine-tuned model Only when asked. A user uploading rows or running an optimize job has not asked to deploy anything, and offering manifests there is noise at best. This skill exists as its own trigger so deployment instructions stay out of sessions that are about data. **Asking about deployment is not asking for files.** "How would I serve this?", "what does vLLM need?", "is my cluster big enough?" are questions: answer them, write nothing. Write manifests only when the user asks for the files, in words like "write the manifests", "add the Deployment", "set this up in my repo". When it is genuinely ambiguous, say what you are about to create and where, and let the user say go. The plugin declares Write, so this gate is the only thing standing between a question and a commit. The precondition, surfaced before anything else: the fine-tune has passed its eval, and for the cluster path there is **at least one GPU node**. ## In the user's own Kubernetes cluster Inference stays inside the user's network. Read `references/deploy-kubernetes.md` and the published [Kubernetes guide](https://docs.runapprentice.com/how-to/deploy-kubernetes), then write the vLLM Deployment and Service into the repo, mirroring the conventions already there (the user's registry, namespace, ingress pattern, label scheme), and state the honest GPU sizing. Never invent cluster names, namespaces, or registries. With no existing manifests visible, ask for one rather than guessing. ## On a Mac Mac-trained MLX adapters are served on the Mac with `mlx_lm.server` ([docs](https://docs.runapprentice.com/how-to/deploy-mlx)). Do not claim an MLX adapter can be served by vLLM elsewhere. That conversion path has no published, verified recipe, and saying otherwise sends a user down a road that dead-ends. ## What deployment does not do Serving a model is not promotion. Activating a model in Apprentice records the promotion in the console and leaves serving unchanged; routing live production traffic to the smaller model is still in development. A user who deploys still decides what calls the endpoint.
Referenced files: 1
apprentice-train4.65 KB
---
name: apprentice-train
description: >
Use when a user is ready to fine-tune a small model and already has
verified rows, or asks what training needs: says "fine-tune", "distill",
"train a small model", "LoRA", or asks how many rows training takes. Also
use for drift on a model already running: whether it still holds quality
and when to retrain. Local MLX training on Apple silicon is the path that
works; hosted training is not shipped yet. For a user still deciding
whether a cheaper model could work, or without a dataset, use apprentice
instead, and delegate recording calls and prompt optimization to
apprentice-capture and serving to apprentice-deploy.
license: MIT
---
# Fine-tune a small model
Training is the second half of the product. The first half, a dataset of verified rows, has
to exist before any of this runs.
## Gold only. This is the part users hit first
Training reads **gold rows only**: rows a human verified. Silver counts for prompt
optimization and never for training.
| | Rows needed |
|---|---|
| Hosted optimize | 20 verified (gold + silver) |
| Local training | enough gold to fill a batch after the held-out split |
| Hosted training | 500 gold, and not yet available (see below) |
A user capturing through an API key accumulates raw traces, and a user uploading rows
accumulates silver. Neither is gold, because recording a verdict needs a signed-in console
session. So the honest first answer to "can I train?" is usually
"not yet, and here is the gap": count the gold rows, name the number missing, and link
`https://runapprentice.com/tasks/<task>/review`.
Do not auto-approve rows to reach 500. Gold means a human checked it. A model grading its own
output is not verification, and every eval downstream inherits the lie.
## Local, on Apple silicon. This is the path that works today
Hosted training is **still being built**, which the
[docs index](https://docs.runapprentice.com/llms.txt) states plainly: the shipped features are
prompt optimization and local training. `client.train(task)` accepts a job, and the server
raises `NotImplementedError` because no GPU trainer is configured. Do not present it as
working, and do not send a user to it.
Local training is real, free, and runs on the user's own Mac via MLX. No Apprentice account is
needed when a CSV is supplied:
```
apprentice train <task> --local --data examples.csv --effort high
```
Or from Python:
```python
client.train_local("duplicate-search", data="examples.csv", effort="high")
```
The model is `Qwen3.5-4B` (Apache-2.0), LoRA fine-tuned.
**`--effort` is a memory profile, not a quality knob.** Every tier trains the same model on
the same examples for the same number of passes; a higher tier uses a bigger batch to finish
sooner and needs more memory. It is auto-detected from the Mac's unified memory. Long prompts
need more memory per example, so drop a tier on an out-of-memory error.
Say this plainly when a user asks whether `--effort low` gives a worse model. It does not.
Run history: `train --local` reaches the console when an API key is set, and never otherwise.
## Reading the result
The number that matters is the held-out score, not the training loss. Held-out rows are the
ones the model never saw, and the eval gate uses gold only.
Two public runs for scale, reproducible in
[apprentice-benchmark](https://github.com/singhabhishekkk/apprentice-benchmark):
- Receipt extraction (OCR text from 200 real scanned receipts, field-level F1, seed 42): a
LoRA fine-tuned Qwen3.5-4B reaches 89.2 on a 60-row held-out split, against GPT-4o-mini at
72.9 plain and 84.2 after GEPA.
- JSON extraction (100 rows, same metric): the fine-tuned 4B reaches 88.9 on a 30-row
held-out split, against 83.1 plain and 85.6 after GEPA.
Those held-out splits are small, so treat them as directional. The user's own run on the
user's own data is the number that decides anything.
Never promise the small model wins. It might not, and the eval exists to find that out.
## Promotion, stated accurately
Activating a model **records** the promotion in the console. It does not move production
traffic: serving is unchanged, and routing live traffic to the smaller model is still in
development. Say this rather than implying a takeover the product does not perform.
## Drift, once a model is live
Capture records the call, feedback records whether it worked, and that feedback score is what
the console's Drift view charts. Without feedback a user sees traffic volume and no quality
signal, which is the common reason "is it still good?" has no answer.
The retrain trigger is a real drop in that score on real traffic, not a calendar date.
## Serving
Once a fine-tune passes its eval, use `apprentice-deploy`.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 12:00 UTC
- Collection status
- Collected
plugins_6a71a0925b0c81919abd1be5add4eabd
Download listing JSON