← AMDCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to AMD
Snapshot Sep 30, 2026 · 23:13 UTC · version 0.2.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "local-ai-use",
"description": "Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API. Use it above all to change that routing persistently, from now on — keep generating pictures locally while chat stays on the cloud; set this workspace up to make images on my own machine — even when the user asks for no image or file in the same breath. Also use it for a single request the user wants done locally, offline, on-device, or kept private: transcribe this recording, make this picture, read this text aloud. Applies in Claude, Cursor, Codex, or any agent harness. Use when the user wants to cut cost or tokens on image, audio, or voice API calls, or to drop DALL-E, hosted Whisper, ElevenLabs, or other paid multimodal APIs; or mentions Lemonade Server, OmniRouter, SD-Turbo, kokoro, Ryzen AI, or NPU/iGPU/dGPU inference. Changes no application source code; do not use it if the user is adding local AI to an app they ship.",
"included_files": [
{
"relative_path": "evals/evals.json",
"size_in_bytes": 2189
},
{
"relative_path": "reference.md",
"size_in_bytes": 9521
},
{
"relative_path": "scripts/setup_local_ai.py",
"size_in_bytes": 24065
},
{
"relative_path": "skill-card.md",
"size_in_bytes": 175
},
{
"relative_path": "templates/local-ai-rule.md",
"size_in_bytes": 5450
}
],
"skill_md_contents": "---\nname: local-ai-use\ndescription: >-\n Makes this agent generate images, transcribe audio, and synthesize speech on\n the user's own machine through a local Lemonade Server instead of a paid cloud\n API. Use it above all to change that routing persistently, from now on — keep\n generating pictures locally while chat stays on the cloud; set this workspace\n up to make images on my own machine — even when the user asks for no image or\n file in the same breath. Also use it for a single request the user wants done\n locally, offline, on-device, or kept private: transcribe this recording, make\n this picture, read this text aloud. Applies in Claude, Cursor, Codex, or any\n agent harness. Use when the user wants to cut cost or tokens on image, audio,\n or voice API calls, or to drop DALL-E, hosted Whisper, ElevenLabs, or other\n paid multimodal APIs; or mentions Lemonade Server, OmniRouter, SD-Turbo,\n kokoro, Ryzen AI, or NPU/iGPU/dGPU inference. Changes no application source\n code; do not use it if the user is adding local AI to an app they ship.\n---\n\n# Local AI Use (route image, TTS, STT through Lemonade)\n\nThis is a **meta-skill**. You run it once. After that, every later request that\nneeds image generation, text-to-speech, or speech-to-text uses the local\n[Lemonade Server](https://lemonade-server.ai) instead of a cloud API. The\nagent's own LLM keeps handling text; only the expensive multimodal calls move\non-device.\n\nThe skill does three things:\n\n1. **Makes sure local Lemonade is installed and running.** If no modern\n `lemonade` CLI is found, the setup script installs the latest version of\n Lemonade on the user's behalf. Modern Lemonade has no `serve` command — the\n Lemonade service (the `lemond` daemon) auto-starts on install and is managed\n by the OS — so the setup script waits for the service and, if it stays down,\n prints the exact OS-specific command to start it (e.g. `sudo systemctl start\n lemond` on Linux).\n2. **Verifies that local Lemonade is reachable.**\n3. **Drops a `Local AI Use` block into the workspace `AGENTS.md`** so the agent\n reads the routing rule on every later turn, in Cursor, Claude Code, Codex,\n Gemini CLI, and any other agent that respects `AGENTS.md`.\n\n> **Requires modern Lemonade (v10.1.0 or newer).** Modern Lemonade unified\n> everything under one `lemonade` CLI (`lemonade status`, `lemonade pull`, ...)\n> driving an always-on `lemond` service. `lemonade` is the only valid CLI, and\n> this skill installs Lemonade only from the [official install\n> paths](https://lemonade-server.ai/docs/guide/install/) listed in Step 1a. If\n> an older `lemonade` is already on the `PATH` it will shadow the modern CLI;\n> uninstall it first (see the removal commands in Step 1a) before running this\n> skill.\n\nModels are **not** downloaded during setup. Each default model is pulled\nlazily, on first use, by the routing rule (e.g. the first image request pulls\nthe image model). This keeps setup fast and avoids gigabytes of downloads the\nuser may never need.\n\n## When to use this skill\n\nUse this skill when **all** of the following are true:\n\n- The user wants local Lemonade. If it is not yet installed, the setup script\n installs the latest version for them automatically.\n- The user accepts the default Lemonade endpoint `http://localhost:13305`.\n- The user wants the change to be **persistent** across future turns and\n agent restarts (the rule is written to disk).\n\nIf the user is instead **embedding** Lemonade as a private subprocess inside\nan app installer, do not use this skill; use `local-ai-app-integration`\ninstead.\n\n## Prerequisites\n\n- **OS:** Windows 11 x64, Ubuntu/Debian x64, or macOS (beta).\n- **Lemonade:** the setup script installs the latest version if missing, using\n `winget` on Windows, the `ppa:lemonade-team/stable` PPA on Ubuntu/Debian, and\n the Homebrew cask on macOS (see Step 1a for the fallbacks). The `lemond`\n service auto-starts after install; the script waits for it rather than\n launching it. On Linux the install needs `sudo`. Pass `--no-install` if the\n user wants to install it themselves instead.\n- **Disk:** ~8 GB free for the three default models (SD-Turbo + Whisper-Tiny\n + kokoro-v1), plus ~0.1 GB for the installer itself. The first image request\n also triggers a ~5 GB pull for `SD-Turbo` if it is not already cached; on\n metered or slow links, consider pulling models eagerly after setup (see\n `lemonade pull` in `reference.md`).\n- **Network:** required for the install download and the first `lemonade pull`\n of each model. After that, every modality runs offline.\n- **Version:** requires v10.1.0 or newer (the unified `lemonade` CLI this\n skill targets). Model IDs and `system-info` fields can change between\n releases; confirm against `lemonade status` and `GET /api/v1/models` on the\n version actually installed rather than assuming this document is current.\n\n## The opinionated path\n\nRun this checklist top to bottom. Track progress against it; do not move on\nuntil each step verifies.\n\n```\n[ ] 1. Ensure Lemonade Server is installed and running (auto-install if missing)\n[ ] 2. Install the routing rule into the workspace AGENTS.md\n```\n\nOn a managed or shared machine where the agent must not run\n`sudo apt-get install`, pass `--no-install` to the setup script and confirm\nLemonade is already installed before continuing.\n\nThe single command that does both steps in one shot is:\n\n```bash\npython scripts/setup_local_ai.py\n```\n\n**Always run this script first — even if Lemonade is already installed and the\nserver is already running, and even before generating a single image.** Writing\nthe routing rule into `AGENTS.md` is what makes this skill complete; skipping it\nbecause \"Lemonade is already up\" leaves the workspace unconfigured for future\nturns. The script is safe to run in that case: it detects the running service,\nskips the install, and just writes the rule.\n\nIt auto-installs the latest version of Lemonade if no modern `lemonade` CLI\nis found, waits for the auto-started `lemond` service, then writes the rule.\nThe script is idempotent: re-running it on a fully configured workspace is a\nno-op apart from a healthcheck. Read the sections below for what to do when\neach step fails.\n\n---\n\n## Step 1: ensure Lemonade Server is installed and running\n\n`scripts/setup_local_ai.py` handles this end to end, but here is what it does\nso you can do it by hand or debug it. Pass `--no-install` when Lemonade is\nalready managed elsewhere and the agent must not attempt a package install.\n\n**1a. Is a modern `lemonade` CLI installed?** Run `lemonade status`. The check\nis by *capability*, not by name: modern Lemonade prints `Server is running...`\nor `Server is not running`. If instead you get an \"invalid choice\" / usage\nerror, the `lemonade` on `PATH` is an old build that predates the unified CLI\n(v10.1.0) — do **not** use it. Remove it, then re-run this skill:\n\n| OS | Uninstall the old build with |\n|---|---|\n| Windows | `winget uninstall -e --id AMD.LemonadeServer`, or Settings > Apps > Installed apps > Lemonade Server > Uninstall |\n| Ubuntu/Debian | `sudo apt remove lemonade-server` |\n| macOS | `brew uninstall --cask lemonade-server`, or delete the installed `Lemonade.app` and its `.pkg` receipt |\n\nNever try to drive or auto-remove it for the user.\n\nIf no `lemonade` is found at all, install the latest version on the user's\nbehalf. Use the package manager first; the download is the fallback when the\npackage manager is absent. Full matrix, including Arch, Fedora, Debian, Snap,\nand Docker, is in the [install docs](https://lemonade-server.ai/docs/guide/install/).\n\n| OS | Install | Fallback |\n|---|---|---|\n| Windows | `winget install -e --id AMD.LemonadeServer` | Download [`lemonade.msi`](https://github.com/lemonade-sdk/lemonade/releases/latest/download/lemonade.msi) and run `msiexec /i lemonade.msi /qn` (silent, per-user, no elevation). |\n| Ubuntu | `sudo add-apt-repository -y ppa:lemonade-team/stable && sudo apt-get update && sudo apt-get install -y lemonade-server` | `sudo snap install lemonade-server` |\n| macOS | `brew install --cask lemonade-server` | Download `Lemonade-<ver>-Darwin.pkg` from the [latest release](https://github.com/lemonade-sdk/lemonade/releases/latest) and run `sudo installer -pkg Lemonade-<ver>-Darwin.pkg -target /`. |\n\nThe Ubuntu apt package is named `lemonade-server`, but the CLI it installs is\n`lemonade`. The browser UI is served at `http://localhost:13305` with no extra\npackage; add `sudo apt install lemonade-desktop` only if the user wants the\ndesktop frontend.\n\nAfter a Windows install the CLI lands in `%LOCALAPPDATA%\\lemonade_server` and\nis added to the *user* PATH (new shells only); the setup script probes that\ndirectory so it works in the same run.\n\n**1b. Is the service running?** Check `lemonade status --json`. The `lemond`\nservice auto-starts on install — there is **no** `lemonade serve` in modern\nLemonade.\n\n| `lemonade status` says | Action |\n|---|---|\n| `Server is running on port 13305` | Continue to Step 2. |\n| `Server is not running` | Wait a few seconds for the auto-started service (the script polls `/api/v1/health`). If it stays down, start it via the OS service manager: `sudo systemctl start lemond` (Linux system install) or `systemctl --user start lemond` (per-user install); `launchctl load /Library/LaunchDaemons/com.lemonade.server.plist` (macOS); the Lemonade tray app or `Start-Service lemond` (Windows). |\n\nOnly if the automatic install genuinely fails (no `apt-get`, no `sudo`,\ndownload blocked) should you stop and point the user at\n<https://lemonade-server.ai/docs/guide/install/>.\n\nThe rest of this skill assumes the endpoint is `http://localhost:13305/api/v1`\nand no API key is required (the system-wide server defaults to no auth on\nloopback). If the user has set `LEMONADE_API_KEY`, the routing rule template\nin `templates/local-ai-rule.md` shows where to add the `Authorization` header.\n\n**1c. Are the backends ready per modality?** Backend health is **per\nmodality**. A working chat or image request does not prove transcription will\nwork: `auto` picks a different backend per modality and can silently fall back\nfor one while having no alternative for another. Before declaring setup\ncomplete, check the actual per-modality state:\n\n```bash\nlemonade backends --all\n```\n\nAny variant the workspace's routing depends on should read `installed`. If the\nonly installed variant for a modality is `rocm`, install the Vulkan variant as\nwell so `auto` has somewhere to fall back to (for example,\n`lemonade backends install whispercpp:vulkan`).\n\n### Default modality models (pulled on first use, not during setup)\n\nSetup does **not** download these. The installed rule pulls each one the first\ntime that modality is requested. They are the smallest models Lemonade offers\nper modality, sized to keep token-and-cost savings real on commodity hardware:\n\n| Modality | Model | Size | Why this default |\n|---|---|---|---|\n| Image generation | `SD-Turbo` | ~5 GB | Single-step generation, runs on CPU and AMD iGPU/dGPU |\n| Text-to-speech | `kokoro-v1` | ~0.3 GB | Only TTS model Lemonade currently supports; CPU-only, low latency |\n| Speech-to-text | `Whisper-Tiny` | ~0.1 GB | Smallest Whisper; fast on CPU. Upgrade to `Whisper-Large-v3-Turbo` if accuracy matters more than latency. |\n\nTo write a different model ID into the rule, pass it to the setup script. For\nexample, to make future image requests use SDXL:\n\n```bash\npython scripts/setup_local_ai.py --image-model SDXL-Turbo\n```\n\nThat model ID is written into the installed `AGENTS.md` rule and pulled on its\nfirst use. The same pattern works for `--tts-model` and `--stt-model`. For\nlarger / higher-quality alternatives (`SDXL-Turbo`, `Flux-2-Klein-4B`,\n`Whisper-Large-v3-Turbo`), see the\n[model picker in reference.md](reference.md#model-picker).\n\n## Step 2: install the routing rule into AGENTS.md\n\nThe rule is a Markdown block stored in [`templates/local-ai-rule.md`](templates/local-ai-rule.md).\nAppend it to the workspace's `AGENTS.md` (create the file if missing). Both\nCursor and Claude Code load `AGENTS.md` automatically on every turn, so the\nagent will see the rule on its next message without any further setup.\n\n`scripts/setup_local_ai.py` does this for you. It bakes the selected endpoint\nand model IDs into the rule, surrounded by stable markers so re-running the\nscript replaces the block in place rather than appending a second copy. The\nmarkers look like:\n\n```\n<!-- BEGIN amd-skills:local-ai-use -->\n...rule...\n<!-- END amd-skills:local-ai-use -->\n```\n\nIf you write the file by hand, keep those exact markers. The script relies\non them for idempotent updates.\n\nIf the user's agent only respects a different convention, mirror the same\nblock to:\n\n- `CLAUDE.md` (Claude Code, project-scoped) or `~/.claude/CLAUDE.md` (global)\n- `.cursor/rules/local-ai-use.mdc` (Cursor user/project rules)\n- `GEMINI.md` (Gemini CLI)\n\nThe rule's content is identical; only the file location changes.\n\n---\n\n## What changes after this skill runs\n\nFrom the next turn onward, the agent reads the rule in `AGENTS.md` on every\nmessage. The rule explicitly tells the agent:\n\n- **For image generation:** call `POST /api/v1/images/generations` on the\n local server. Do **not** call any cloud image API and do **not** use the\n built-in `GenerateImage` tool (that path bills tokens to the cloud\n provider).\n- **For text-to-speech:** call `POST /api/v1/audio/speech`. Do **not** call\n cloud TTS providers (OpenAI TTS, ElevenLabs, etc.).\n- **For speech-to-text:** call `POST /api/v1/audio/transcriptions`. Do\n **not** call cloud transcription providers.\n- **Fallback:** only fall back to a cloud API after one local attempt has\n failed *and* the user has been told the local call failed. Never silently;\n the whole point of this skill is to keep cost predictable. For\n speech-to-text, the disclosure must also say the transcript came from a\n different engine, since mixed-engine transcripts should not be compared or\n deduplicated.\n\nThe agent's own text reasoning continues to use whatever LLM Cursor / Claude\nCode / Codex is configured with. This skill does not redirect chat tokens;\nit only redirects the multimodal calls that would otherwise leave the\nmachine.\n\n## Troubleshooting cheatsheet\n\n| Symptom | Cause | Recovery |\n|---|---|---|\n| `lemonade: command not found` | CLI not installed | Re-run `python scripts/setup_local_ai.py` (auto-installs the latest version). If it just installed on Windows, open a new shell so the user PATH refreshes, or the script will find it under `%LOCALAPPDATA%\\lemonade_server`. |\n| `status` gives an \"invalid choice\" / usage error | An old `lemonade` (pre-v10.1.0) is shadowing the modern CLI | Uninstall it (see the Step 1a table: `winget uninstall -e --id AMD.LemonadeServer` / `sudo apt remove lemonade-server` / `brew uninstall --cask lemonade-server`), then re-run the setup script. |\n| `Server is not running` | `lemond` service stopped | Start it via the OS service manager — `sudo systemctl start lemond` / `systemctl --user start lemond` (Linux), `launchctl load /Library/LaunchDaemons/com.lemonade.server.plist` (macOS), or the tray app / `Start-Service lemond` (Windows). There is no `lemonade serve`. |\n| `POST /v1/images/generations` returns 404 model not found | Image model not downloaded | `lemonade pull SD-Turbo` and retry. |\n| `lemonade pull` keeps printing `Progress: NN%` but never finishes | Download target is a bad path (out of space, no write permission, quota, read-only mount). The write error may surface only in the server log while the console keeps showing progress | Check the target and free space first: `GET /api/v1/system-info` reports `models_dir` and `model_storage.free_bytes`. If a pull stalls, read the recent lines of the server log (typically `lemonade-server.log` in the OS temp dir) for the real error (e.g. a download/write failure like `CURL code 23`, or an out-of-space message), then point the download at a writable disk with room. |\n| Image generation is slow on CPU (~4–5 min) | sd-cpp on CPU backend | Install the GPU backend on supported AMD hardware: `lemonade backends install sd-cpp:rocm`. |\n| Still slow after installing the GPU backend | The backend is installed but not actually engaged; the runtime fell back to CPU silently | An `installed` state in `system-info` and a successful `rocminfo` both still permit a silent CPU fallback. Check real GPU utilisation (`gpu_busy_percent`) during a request, and confirm the host's GPU driver stack rather than re-installing the backend. |\n| `POST /v1/audio/transcriptions` returns 400 unsupported format | Input is not 16 kHz mono WAV | Re-encode with `ffmpeg -i in.* -ar 16000 -ac 1 out.wav`. |\n| `POST /v1/audio/speech` returns 404 | TTS model not downloaded | `lemonade pull kokoro-v1`. |\n| 401 Unauthorized on every request | User has set `LEMONADE_API_KEY` | Add `Authorization: Bearer $LEMONADE_API_KEY` to every request and to the rule block. |\n\n## Verification checklist\n\nMark this skill complete only when **all** of the following are true:\n\n- [ ] `lemonade status --json` reports the server running on port 13305.\n- [ ] The workspace `AGENTS.md` contains the\n `amd-skills:local-ai-use` block. This is required even when Lemonade was\n already installed and running — generating an image alone does not\n complete the skill.\n- [ ] On a follow-up turn, asking the agent to \"generate an image of X\"\n causes it to POST to `http://localhost:13305/api/v1/images/generations`\n (pulling the model on first use) rather than calling a cloud tool.\n- [ ] `lemonade backends --all` shows `installed` for every backend variant\n this workspace's routing depends on (see Step 1c). Do not treat a working\n image or chat path as proof that transcription will work.\n\nIf any box is unchecked, the user is still paying cloud cost for at least\none modality, or a routed modality may fail silently on first use.\n\n---\n\n## Reference\n\nFor the full model picker, alternate-quality options, the complete endpoint\nreference, the API-key flow, and the OmniRouter tool definitions you can\nhand to an agent's tool-calling loop, see [reference.md](reference.md).\n"
}SHA-256: aa78467afefd0585a5904ee8d26dc7b156f174d5699e4cbda55c3ea046b8530a