{"id":16987,"plugin_id":"plugins_6a71a0925b0c81919abd1be5add4eabd","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:13:55.435Z","digest":"ffba5b1beaba2cc389af5965b8ba2276852ded99486761655c3c8d357d318bb7","against":null,"payload":{"description":"Use when a user wants real LLM calls recorded into an Apprentice dataset or a prompt optimized: \"capture my calls\", \"record traces\", \"log these to Apprentice\", \"optimize this prompt\", or after the apprentice skill flagged a repeatable call and the user agreed. Wires the capture line into the code that makes the calls, uploads rows, runs optimize, and returns the console link for verifying rows or sending them to an expert. Delegate fine-tuning to apprentice-train and serving to apprentice-deploy.","included_files":[{"relative_path":"references/use-optimized-prompt.md","size_in_bytes":837}],"name":"apprentice-capture","skill_md_contents":"---\nname: apprentice-capture\ndescription: >\n  Use when a user wants real LLM calls recorded into an Apprentice dataset\n  or a prompt optimized: \"capture my calls\", \"record traces\", \"log these to\n  Apprentice\", \"optimize this prompt\", or after the apprentice skill flagged\n  a repeatable call and the user agreed. Wires the capture line into the\n  code that makes the calls, uploads rows, runs optimize, and returns the\n  console link for verifying rows or sending them to an expert. Delegate\n  fine-tuning to apprentice-train and serving to apprentice-deploy.\nlicense: MIT\n---\n\n# Capture calls and optimize the prompt\n\nDo the job. A plan handed back is not the job.\n\n## Pick the path. Do not ask which one\n\n| What the user said | What to use |\n|---|---|\n| Nothing about accounts, and it is a Python app | **SDK.** Wire capture, hand back the console link |\n| \"I don't want to sign in\", \"keep it local\", \"no account\" | **CLI** `--local`, the user's own OpenAI key |\n| \"Just optimize this prompt\", and rows already exist | **SDK.** Upload, then run |\n| Already in the console | Send a deep link back to it |\n\nAsk only when the choice changes what the user gets. Say which path in one line.\n\nThe SDK is the default because it returns real values instead of text to parse. The CLI earns\nits place only for the no-account case.\n\nDo not mix the two in one piece of work. A session that used the SDK for uploads and the CLI\nfor status checks left the user unable to tell which interface had done what.\n\n## Wire capture in, do not just describe it\n\nOne line beside the existing model call. It is fail-open by design: it returns `None` rather\nthan raising, so a capture outage cannot take down a user's endpoint.\n\n```python\nfrom runapprentice import Apprentice\n\nclient = Apprentice(api_key=os.environ[\"APPRENTICE_API_KEY\"])\ntrace_id = client.capture(task=\"duplicate-search\", input=question, output=answer)\n```\n\nThen say which file changed and how to remove it. One line to undo is what makes doing it\nsafe: a user who dislikes it reverts in seconds, a user who was only offered it has nothing.\n\n**Never run a real flow from a throwaway script and then delete it.** The dataset stops\ngrowing the moment that script is gone, and nothing in the repo records it happened. A real\nsession did exactly this: eight `/tmp` scripts, zero repo changes, and a dataset frozen at six\nrows because \"you had not asked for ongoing capture\". True, and useless.\n\n### Framework-specific capture\n\n- LangChain: `ApprenticeCallback` captures calls and simple retriever context.\n  [Guide](https://docs.runapprentice.com/how-to/capture-langchain). Use manual `capture(...)`\n  when there are several retrievers or custom context formatting.\n- Raw OpenAI clients, Chat Completions and Responses:\n  [guide](https://docs.runapprentice.com/how-to/capture-openai).\n- Full method list: [Python SDK reference](https://docs.runapprentice.com/reference/python-sdk).\n\nTwo API details worth getting right, both from the docs' rules for coding agents: upload with\n`client.datasets.upload(...)`, since there is no `ingest()` method, and pass structured\n`inputs={...}` for a multi-field or templated task rather than one rendered prompt string. For\nRAG, `inputs={\"question\": question, \"context\": exact_context}`, where the context is exactly\nwhat the model saw.\n\n## Feedback is what makes drift measurable\n\nCapture records the call. Feedback records whether it worked, and that score is what the\nconsole's Drift view charts and what decides when a retrain is worth doing.\n\n```python\nif trace_id:                              # None when capture failed, by design\n    client.feedback(trace_id, good=True)  # or good=False, or score=0.4\n```\n\n**Never manufacture it.** Do not add a second LLM call to grade the first, and do not infer\n\"looked fine\" from confidence or length. A guessed score retrains the model on a lie. With no\nreal signal, wire nothing: captured rows still become gold when a human verifies them.\n\n## Optimize\n\n```python\njob = client.optimize(\"duplicate-search\").wait()\n```\n\nHosted optimize needs **20 verified rows**, counting gold plus silver. Below that it refuses,\nand the fix is verification, not more uploading.\n\nFor the no-account case, with the user's own OpenAI key:\n\n```\napprentice optimize <task> --local --data examples.csv\n```\n\nLocal optimize is scored as JSON extraction only, so every output must be a JSON object or\narray; the CLI refuses other data before spending anything. Note also what the console shows:\n`optimize --local` is never recorded there, so that run stays on the machine that ran it.\n\n## Hand back the link, every time\n\n| Page | URL |\n|---|---|\n| Rows and tiers | `https://runapprentice.com/tasks/<task>/dataset` |\n| Review queue | `https://runapprentice.com/tasks/<task>/review` |\n| Runs | `https://runapprentice.com/tasks/<task>/runs` |\n| Send to an expert | `https://runapprentice.com/tasks/<task>/collaborators` |\n\nSay what is waiting there, not just the address: \"6 rows uploaded as silver, verify them at\n`https://runapprentice.com/tasks/duplicate-search/review`\".\n\n## An API key cannot make rows gold, and the two paths do not land in the same place\n\nUploaded rows land as **silver**. Captured traces land as **raw**, which is one step further\nback: raw counts for nothing until a human reviews it, while silver already counts toward\noptimize. Recording a verdict needs a signed-in console session either way, so a key-only\nworkflow can never promote either kind.\n\nThat matters because the two features count differently: **optimize uses gold plus silver,\ntraining uses gold only**. So a user who uploads rows can optimize immediately and can never\ntrain; a user who only captures cannot even optimize until the traces are reviewed. Say which\none applies at the moment the rows land, not when the user hits a threshold, and link the\nreview page so it is one click to fix.\n\nDo not auto-approve rows to gold to clear a threshold. Gold means a human checked it, and a\nmodel grading its own output is not that.\n\n## After a run finishes\n\nRead `references/use-optimized-prompt.md` when wiring the optimized prompt into the user's\ncode. The artifact is not a template and `.format()` on it fails silently in a way that ships\na prompt which never sees the user's data.\n\n## Sandbox\n\n`Operation not permitted` on a network call means the sandbox denied it, so the request never\nleft the machine. A bare connection error is weaker evidence and has other causes (DNS, TLS, a\nproxy, a firewall, an outage), so name the sandbox as one possibility rather than the answer.\nClaude Code and Codex sandboxes deny network by default; tell the user what to enable rather\nthan retrying or working around it.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}