← Plugin catalog
Developer Tools

AMD

AMD v0.2.0

AMD's verified Agent Skills in one plugin: route image/audio through local AI on Ryzen AI, serve LLMs on AMD Instinct GPUs with vLLM, optimize inference throughput with Hyperloom, and analyze GPU kernel and PyTorch trace performance.

Language: English · Automatically detected from descriptions.

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
MIT
Package author
AMD
Keywords
amd, rocm, hip, ryzen-ai, vllm, lemonade, local-ai

Declared capabilities

  • Read
  • Write

Package observed Sep 30, 2026.

Files & skills

File archives

Plugin package144 files · 610 KBBrowse files →
Skill instructions
hyperloom-workload-optimizer20.4 KB

View saved version →

---
name: hyperloom-workload-optimizer
description: >-
  Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs
  and reports a validated gain, using the Hyperloom multi-agent optimizer. Given a
  model, framework, workload (TP/EP, concurrency, ISL/OSL, precision), an objective
  and a time budget, it explores per-workload which levers to pull (serving/config
  parameters and env, framework enablement and source patches, and hot GPU-kernel
  rewrites), benchmarks each candidate, and returns the optimization stack that
  produced the gain. Use when the user wants to make a model serve faster, raise
  tokens/sec or throughput, optimize or tune vLLM or SGLang on MI300X/MI325X/MI355X,
  run Hyperloom, run the kernel-agent, quantize-then-optimize with Quark, set up
  Hyperloom from scratch, or resume a Hyperloom session. Do not use to stand up a
  server for plain serving, diagnose a broken ROCm install, or run a one-off
  kernel/benchmark or trace analysis without the optimization loop.
---

<!--
Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.

See LICENSE for license information.
-->

# Hyperloom Workload Optimizer

You are the catalog entry point for Hyperloom optimization on AMD Instinct GPUs.
Bootstrap the workspace, prepare the runtime environment, collect workload
parameters, then install, launch, and monitor the optimizer. This skill owns the
orchestration and the launcher gates; environment prep and workload intake are
delegated to the skills the Hyperloom wheel installs, and
`@${HYPERLOOM_SKILL_PATH}` (`inference_optimizer`) is the execution baseline.

Do not manually optimize inside chat unless debugging.

## Prerequisites

- AMD Instinct GPU host (MI300X / MI325X / MI355X) with ROCm
- `/dev/kfd` and `/dev/dri` present; `amd-smi` or `rocm-smi` works
- Python 3.10+ and network access to install the Hyperloom wheel
- Anthropic (or compatible) LLM credentials for agent backends
- A dedicated agent workspace directory

Every command in this skill runs on that GPU host. Confirm the shell you are in
is on it before Phase 0, so a bootstrap does not land on a machine with no GPU.

The Hyperloom **runtime** ships via `pip install` of the published wheel.

## What Hyperloom runs

The CLI starts a Python Coordinator that coordinates:

- **Orchestration** — baseline, explore, specialist, integrate_patch, sweep
- **Kernel** — trace_analyze, run_optimization, integrate
- **Critic** — proposal review (default `--critic-agent`)
- **Robustness** — health monitoring and RCA (default `--robustness-agent`)

State lives under a **session directory** per run; run-state root is
`$USER_DATA_PATH` (default `/workspace/hyperloom`), independent of the install
directory (`INSTALL_DIR`, where the wheel and `.env` live) and may point to
shared storage. Layout: `$USER_DATA_PATH/runtime/` (install.sh outputs,
`kernel-agent.env.sh`), `logs/`, and `<model_basename>/<UTC_ts>/` per session
holding `manifest.json`, `state.json`, `runs/`, `reports/`, `optimizer_runs/`.

## Workflow overview

Match `hyperloom-custom-advanced` section order — do **not** ask workload
questions while writing `.env` or during `/hyperloom-setup`.

- **Phase 0 Bootstrap** — `pip install`, `/hyperloom-setup` → `.env` (credentials + run mode only)
- **Phase 1 Environment** — custom-advanced §Setup Configuration (baremetal: confirm host; docker: start container + setup inside, contract in [setup.md](setup.md))
- **Phase 2 Workload intake** — custom-advanced §Advanced Configuration → Model Resolution → show launch plan → user confirms
- **Phase 3 Execute** — install.sh → preflight → launch → monitor → report

Load `hyperloom-custom-advanced` at Phase 1 and follow its sections in order
(discovery: `.cursor/` / `.claude/` / `.agents/skills/hyperloom-custom-advanced/SKILL.md`).
If it is not on disk, stop and tell the user to restart the agent so the newly
installed skills are picked up — do not improvise the environment or workload
sections from memory, since the wheel is the source of truth for both.
For deeper optimizer behavior read `@${HYPERLOOM_SKILL_PATH}` (`inference_optimizer`);
Iron Rules + CLI reference: [reference.md](reference.md).

## Iron Rules (launcher gates)

Run order is always **IR-2 → IR-1 → launch**. Full text in [reference.md](reference.md).

- **IR-1 — GPU unoccupied.** Before every `optimize` (fresh or `--resume`), every
  visible GPU must have zero foreign serving PIDs (`sglang.launch_server` /
  `vllm.entrypoints` / `Magpie`) and ≲ 500 MiB VRAM in use.
- **IR-2 — install.sh before launch.** Run `install.sh` and source
  `kernel-agent.env.sh` in the **same shell** that spawns `optimize`.
- **Resume carve-out:** `--resume` may skip install only when `install.sh` exited
  0 earlier in the same shell, `kernel-agent.env.sh` is still sourced, and the
  session's `manifest.json` exists. Any failure → re-run `install.sh`.

## Phase discipline (do not skip)

One phase at a time. Each phase asks only its own questions, waits for the
user's answers, completes its exit condition, then moves on. Never batch
questions from different phases into one prompt. In particular, never ask
workload questions (model, framework, TP/EP, precision, ISL/OSL, hours…) during
Phase 0 or Phase 1 — those belong to Phase 2 only.

## Phase 0 — Bootstrap

Skip completed steps (idempotent). Ask only about the install directory and
credentials/run mode here. Do not ask about the model or workload yet.

### Confirm the install directory

The wheel installs into a target directory with `pip install --target <dir>`,
which also holds `.env` and runtime artifacts. Do not silently use the current
directory. Show the resolved current directory (`pwd`) and confirm it with the
user, or let them choose another dedicated path. Wait for the answer, then `cd`
into the chosen directory before installing.

### Install the Hyperloom wheel

Skip when `hyperloom/` (wheel) or `src/hyperloom/` (source) already exists in the
confirmed directory.

The runtime is published to PyPI as `hyperloom-inference-optimizer`. List the
releases, tell the user the newest one, and ask whether to install it or a
version they name.

List with `--pre` so prereleases are visible, and install an exact `==` version
so a later bootstrap installs the same runtime.

```bash
cd "$INSTALL_DIR"   # the directory confirmed above
pip index versions hyperloom-inference-optimizer --pre
pip install hyperloom-inference-optimizer==<version the user approved> --target .
```

Confirm `hyperloom/inference_optimizer/assets/install.sh` exists. Restart the
agent if wheel skills are not visible.

### Credentials and run mode

Run `/hyperloom-setup` (installed to `.cursor/skills/hyperloom-setup/`). It
writes `.env`, sets `USER_DATA_PATH`, `HYPERLOOM_RUN_MODE`, and
`HYPERLOOM_SKILL_PATH`, and on bare metal runs `install_baremetal.sh`.

**Phase 0 is done when all hold:**

- `hyperloom/inference_optimizer/assets/install.sh` exists
- `.env` exists with non-placeholder LLM secrets
- `USER_DATA_PATH`, `HYPERLOOM_RUN_MODE`, and `HYPERLOOM_SKILL_PATH` are set

More bootstrap detail: [setup.md](setup.md).

## Phase 1 — Environment prep

Load `hyperloom-custom-advanced` and follow its **Setup Configuration**
section only.

**Baremetal (`HYPERLOOM_RUN_MODE=baremetal`):** confirm `install_baremetal.sh`
finished and the serving framework from setup is importable. Do not ask workload
questions yet.

**Docker (`HYPERLOOM_RUN_MODE=docker`):** image choice, `docker run`, and the
in-container setup are owned entirely by custom-advanced Setup Configuration —
follow it, do not restate its commands or flags here. Do not ask workload
questions until the container is up and in-container setup succeeded, and never
run `optimize` on the host.

Phase 1 is done when the target environment (host or container) is ready.

## Phase 2 — Workload intake

Enter only after Phase 0 and Phase 1 exit conditions hold. This is the first and
only phase that asks workload questions.

Now follow custom-advanced **Advanced Configuration**, **Default Values**, and
**Model Resolution**. Use the agent's structured question UI when available.
Never copy API keys into chat output.

| Field | CLI flag | Default | Notes |
|---|---|---|---|
| Model path | `--model` | required | Local dir with `config.json`, or HF cache |
| Framework | `--framework` | `sglang` | or `vllm`; prefer `.env` `FRAMEWORK` when set |
| TP / EP | `--tp` / `--ep` | `1` / `1` | tensor / expert parallel |
| CONC | `--conc` | `64` | client concurrency |
| ISL / OSL | `--isl` / `--osl` | `1024` / `1024` | input / output seq lengths |
| PRECISION | `--precision` | `bf16` | match checkpoint; `fp8` for FP8 models |
| MAX_HOURS | `--max-hours` | CLI `2.0` | offer `3` (quick) or `12` (full); see below |
| TARGET_GAIN | `--target-gain` | `30` | desired % gain |

**Optional:** `--no-explore`, `--no-enable-conc-sweep`, `--gpu-type`,
`--server-args`, `--compare-against-gpu`, `--quantize` prelude.

Infer `PRECISION` from the model name when obvious (e.g. an `FP8` model implies
`--precision fp8`) and confirm it — do not silently keep the `bf16` default.

### Budget and flags — offer these three

Offer all three and let the user pick one. The flags in each are a set: pass them
together, and do not ask for a budget and then ask separately which phases to run.
The two demos take the workload and flags of the Hyperloom demo skill of the same
budget — treat those as given and skip the table above. The user may name their
own model instead of the demo's; for the 3-hour demo keep it at 8B or below.
Confirm everything in the launch plan. Only **Custom** collects workload answers.

**1. 3-hour demo** (`hyperloom-qwen3-8b-3h`) — `Qwen/Qwen3-8B` unless the user
names another 8B-or-smaller model, TP=1, CONC=64, ISL=OSL=1024,
`--precision bf16`, serving and config parameters only, no kernel rewrites.
Resolve the model per custom-advanced Model Resolution; download it from Hugging
Face when it is not already local. Match `--precision` to the chosen checkpoint.
Expect a modest validated gain, or an honest 0% when the workload has no
parameter headroom.

```text
--max-hours 3 --precision bf16
--no-framework-agent --no-kernel --no-enable-conc-sweep --no-enable-roofline
--max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01
--explore-force-exit-budget-pct 0.01 --explore-force-exit-hours-remaining 0.05
```

**2. 12-hour demo** (`hyperloom-qwen3-14b-fp8-12h`) — `Qwen/Qwen3-14B-FP8`
unless the user names another model, TP=1, CONC=64, ISL=OSL=1024,
`--precision fp8` matched to the chosen checkpoint, every lever with kernel
rewrites included. The kernel agent needs room to profile, rewrite and
revalidate, which is where the larger gains come from.

```text
--max-hours 12 --precision fp8
--max-minutes-framework-pct 0.01 --max-minutes-explore-pct 0.42
--max-minutes-kernel-pct 0.42
```

**3. Custom** — the user brings their own model or workload instead of taking a
demo. Walk through the fields in the table above and the phase toggles, one
question at a time, and derive the flags from the answers rather than asking for
flags. Whichever levers they pick, a budget of 3 hours or less keeps the 3-hour
demo's flag set.
Optional flags come from the list above; show the full flag list in the launch
plan either way.

### Confirmation gate (required before Phase 3)

The Coordinator has no in-loop `setup` / `classify` — a value not asked here is
silently lost to its default. Before running any Phase 3 command, present the
full launch plan (including defaulted fields) and get explicit user confirmation.

Print the plan in the reply body as this aligned block:

```text
Launch plan — please confirm:
  MODEL_PATH    /wekafs/models/Qwen3-14B-FP8
  FRAMEWORK     vllm
  TP=1  EP=1  CONC=64
  ISL=1024  OSL=1024
  PRECISION=fp8
  MAX_HOURS=3     TARGET_GAIN=20%
  profile       3-hour demo — no kernel, no framework agent, no roofline
  flags         --no-framework-agent --no-kernel --no-enable-conc-sweep
                --no-enable-roofline
                --max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01
                --explore-force-exit-budget-pct 0.01
                --explore-force-exit-hours-remaining 0.05
  RUN_MODE      baremetal
```

Never put the plan inside the confirmation prompt itself. A prompt renders as one
wrapped paragraph, which collapses the alignment above into an unreadable blob the
user has to search for `MAX_HOURS` in. Keep the prompt to a single short question
such as `Approve this launch plan?`, and if you offer a "change something" option,
name the field to change rather than making the user retype it as free text.

Do **not** run `install.sh` or launch `optimize` until the user approves this plan.

### Persist the plan (required — shells do not share exports)

Agent shells do not persist exports between calls, so write the confirmed values
to `$RUN_DIR/workload.env` right after approval. Every Phase 3 block sources it;
without this, launch silently falls back to `${TP:-1}` / `${CONC:-64}` defaults
and `--model ""`. Fill each value from the approved plan.

```bash
export USER_DATA_PATH="${USER_DATA_PATH:?run /hyperloom-setup first}"
export RUN_DIR="${USER_DATA_PATH}/optimizer_runs"
mkdir -p "$RUN_DIR"
# Quoted heredoc (<<'EOF'): values are written literally, so a MODEL_PATH with
# spaces, $, or $(...) is not expanded or executed. Edit each value to the plan.
cat > "$RUN_DIR/workload.env" <<'EOF'
export MODEL_PATH=/wekafs/models/Qwen3-14B-FP8
export FRAMEWORK=vllm
export TP=1
export EP=1
export CONC=64
export ISL=1024
export OSL=1024
export PRECISION=fp8
export MAX_HOURS=3
export TARGET_GAIN=20
# The whole flag set for the approved profile, space-separated. The 3-hour
# demo is shown; a 12-hour run swaps in its own set.
export OPT_FLAGS="--no-framework-agent --no-kernel --no-enable-conc-sweep --no-enable-roofline --max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01 --explore-force-exit-budget-pct 0.01 --explore-force-exit-hours-remaining 0.05"
EOF
```

## Phase 3 — Install (IR-2)

`INSTALL_DIR` is the directory confirmed in Phase 0, the one holding `hyperloom/`
and `.env`. Every Phase 3 block below rebuilds it from the current directory, so
run them from there; the check refuses a directory that is not it.

Resolve paths for wheel or source layout:

```bash
export INSTALL_DIR="$(pwd -P)"
[ -d "${INSTALL_DIR}/hyperloom" ] || [ -d "${INSTALL_DIR}/src/hyperloom" ] || {
  echo "ERROR: ${INSTALL_DIR} holds no hyperloom/ -- cd to the Phase 0 install directory" >&2; exit 1; }
set -a; . "${INSTALL_DIR}/.env"; set +a
export USER_DATA_PATH="${USER_DATA_PATH:?USER_DATA_PATH missing}"
. "${USER_DATA_PATH}/optimizer_runs/workload.env"   # confirmed Phase 2 values
export PYTHONPATH="${INSTALL_DIR}:${PYTHONPATH:-}"
ulimit -Sn 65536 || true

INSTALL_SH="${INSTALL_DIR}/hyperloom/inference_optimizer/assets/install.sh"
[ -f "$INSTALL_SH" ] || INSTALL_SH="${INSTALL_DIR}/src/hyperloom/inference_optimizer/assets/install.sh"

bash "$INSTALL_SH"
. "${KERNEL_AGENT_ENV:-${USER_DATA_PATH}/runtime/kernel-agent.env.sh}"
export PYTHONPATH="${INSTALL_DIR}:${PYTHONPATH:-}"
```

In Docker mode, run this inside the container.

## Phase 3 — Preflight (IR-1)

`install.sh` exports `$PYTHON`; the fallback below covers agent sandboxes that do
not persist exports between shell calls.

```bash
export SKILL_DIR="${SKILL_DIR:?absolute path of the directory holding this SKILL.md}"
. "${USER_DATA_PATH}/optimizer_runs/workload.env"   # confirmed Phase 2 values
export PYTHON="${PYTHON:-$(command -v python3)}"
"$PYTHON" "${SKILL_DIR}/scripts/preflight.py"
```

The gate exits non-zero — do not launch — when `MODEL_PATH` is missing or has no
`config.json`, torch sees no GPU, a foreign serving process still holds a card,
or any GPU holds more than `IR1_VRAM_LIMIT_MIB` (default 500) MiB.

It also blocks when VRAM cannot be read at all: no `amd-smi`/`rocm-smi` on
`PATH`, a probe that exits non-zero, or output it cannot parse. An unreadable
probe cannot rule out a busy GPU, and a foreign process holding VRAM under a
different name would slip through. Confirm the GPUs are idle by hand before
re-running with `IR1_ALLOW_UNVERIFIED_VRAM=1`.

Never print API keys or tokens. `scripts/tests/test_preflight.py` covers the
probe shapes this gate must reject.

## Phase 3 — Launch

After IR-2 and IR-1 pass, launch. `setsid nohup` is required for runs longer than
5 minutes, so the run outlives the agent shell.

```bash
export INSTALL_DIR="$(pwd -P)"
export SKILL_DIR="${SKILL_DIR:?absolute path of the directory holding this SKILL.md}"
bash "${SKILL_DIR}/scripts/launch.sh"
```

Every workload value comes from the confirmed `workload.env`; the script has no
`${VAR:-default}` fallbacks, so a missing value fails loudly instead of launching
a different config. Put any optional Phase 2 flags (`--no-kernel`, `--no-explore`,
`--gpu-type`, `--model-class`, `--server-args`, `--compare-against-gpu`,
`--quantize`, phase budget flags) into `OPT_FLAGS` in `workload.env`. `OPT_FLAGS`
is word-split, so quote any flag value that contains spaces, e.g.
`export OPT_FLAGS='--server-args "--foo bar"'`.

### Launch health check (30 s after start)

Required after every launch and resume. The PID recorded at launch is the
**setsid wrapper**, which exits immediately — it is NOT the optimizer. This reads
the real `.pid` and `.session_dir` from the launch-info JSON, rewrites the PID
file so the monitor watches the right process, and records both in
`$RUN_DIR/last_launch.env` for the later phases.

```bash
export INSTALL_DIR="$(pwd -P)"
export SKILL_DIR="${SKILL_DIR:?absolute path of the directory holding this SKILL.md}"
bash "${SKILL_DIR}/scripts/launch_health.sh"
```

It exits non-zero when the launch-info JSON never appeared, no optimizer process
can be found, or `session_dir` is still unset — inspect the reported run log in
those cases. Never guess `session_dir` from a timestamp; concurrent sessions
share `USER_DATA_PATH`.

## Phase 3 — Monitor

Poll at most every 5 minutes unless debugging a startup failure. Use the state
reader the wheel ships rather than parsing `state.json` by hand — it also prints
the recent lifecycle events.

```bash
export INSTALL_DIR="$(pwd -P)"
. "${USER_DATA_PATH}/optimizer_runs/last_launch.env"   # SESSION_DIR from launch
STATE_TOOL="${INSTALL_DIR}/hyperloom/inference_optimizer/tools/read_optimizer_state.py"
[ -f "$STATE_TOOL" ] || STATE_TOOL="${INSTALL_DIR}/src/hyperloom/inference_optimizer/tools/read_optimizer_state.py"
"${PYTHON:-python3}" "$STATE_TOOL" "$SESSION_DIR"
```

For recent action counts grouped by category, the wheel also ships
`tools/event_counts.py`, invoked the same way.

Report session id + log path, `baseline_tput` / `current_best` /
`cumulative_gain`, explore accepted/rejected, last kernel opt (correctness,
speedup, KEEP/REVERT), and process-alive vs `stop_reason`. See
[reference.md](reference.md) Report fields.

## Resume

Resume runs in a fresh shell. Re-run the IR-2 and IR-1 gates first, exactly as for
a fresh launch — the script does not re-check them.

```bash
export INSTALL_DIR="$(pwd -P)"
export SKILL_DIR="${SKILL_DIR:?absolute path of the directory holding this SKILL.md}"
bash "${SKILL_DIR}/scripts/resume.sh"
bash "${SKILL_DIR}/scripts/launch_health.sh"
```

It resumes the session recorded in `last_launch.env` and always passes
`--resume-from` explicitly, because a bare `--resume` auto-picks the newest
session and can target the wrong run. Resume writes its own log
(`run_resume-*.log`) so the original run log is preserved. Reuse the IR-2
carve-out rules; re-run `install.sh` if the shell or env changed.

| `stop_reason` | Action |
|---|---|
| `time_exhausted` | `--resume` same session |
| `no_more_leverage` | stop; resume only if user changes strategy |
| `policy_loop` | inspect `policy_denial_history`; clear stale prunes |

## Expected optimizer flow

1. Establish `baseline_tput`.
2. Coordinator runs roofline/profile analysis after baseline.
3. `explore` tests serving parameters incrementally.
4. Kernel-agent runs on hot paths with compile + correctness evidence.
5. `sweep` validates concurrency around the best candidate.
6. Final report under `$SESSION_DIR/reports/`.

## When to defer

- **Plain serving only** — use `serving-llms-on-instinct`.
- **ROCm driver broken** — diagnose the ROCm stack first (e.g. a `rocm-doctor`
  skill if published); do not start the optimizer on a broken driver.
- **Edge cases** — read `@${HYPERLOOM_SKILL_PATH}` for multi-node, atom
  framework (IR-8), critic/robustness backends, cache topology, and the
  full failure matrix.

## Further reading

- Bootstrap detail: [setup.md](setup.md)
- Iron Rules + CLI reference: [reference.md](reference.md)
- Authoritative runtime skill: `hyperloom/inference_optimizer/SKILL.md`

Referenced files: 12

local-ai-app-integration23.6 KB

View saved version →

---
name: local-ai-app-integration
description: >-
  Integrates local AI capabilities into applications using Embeddable Lemonade.
  Use when the user wants to add local AI, offline AI, private AI, on-device AI,
  a local LLM, local chat, embeddings, image generation, speech-to-text, or
  text-to-speech to an existing app; replace or supplement OpenAI, Anthropic, Ollama, or
  other cloud AI APIs with a local backend; only use to convert user apps. Do not use when
  the user just wants the agent itself to generate images, transcribe, or speak locally in
  the current workspace, even to cut their own API bill.
---

# Local AI App Integration (Embeddable Lemonade)

Add a local AI mode to an existing app that already talks to a cloud AI API
(OpenAI, Anthropic, or Ollama-compatible). The app launches `lemond`, the
Embeddable Lemonade binary, as a private subprocess and the existing client
talks to it on `http://localhost:PORT/api/v1`. The user gets local, private,
hardware-optimized inference (CPU, AMD iGPU/dGPU, XDNA2 NPU) with no separate
install.

**What you'll end up with:** one new launcher module (~30 lines), three mandatory changes to the existing HTTP client (`base_url`, `api_key`, and a 120-second HTTP timeout), one vendored binary under `vendor/lemonade/`.

## When this skill is the right tool

Use this skill when **all** of the following are true:

- The app already calls a cloud AI service over HTTP (OpenAI Chat Completions,
  Anthropic Messages, or Ollama).
- The user wants that AI to run on the end-user's PC, with the AI engine
  bundled into the app, not as a separate user install.
- The target platform is Windows x64 or Linux x64 (macOS embeddable is in beta).

If the user instead wants a **system-wide** Lemonade Server (one install,
shared across apps), do not use this skill; point them at
`https://lemonade-server.ai/install_options.html` and the standard OpenAI base
URL `http://localhost:13305/api/v1`.

## The opinionated path

This skill follows one fixed sequence. Do not deviate without a stated reason.

```
[ ] 1. Survey the app's current AI integration
[ ] 2. Pick a model + backend profile
[ ] 3. Place Embeddable Lemonade in the app's tree (full package, not just the binary)
[ ] 4. Add a `lemond` launcher (subprocess + API key + port + per-stage logging)
[ ] 5. Re-point the existing client at lemond (base_url, api_key, 120s timeout — all three required)
[ ] 6. Wait for /api/v1/health, install backend, then PULL the model before first use
[ ] 7. Wire shutdown and error recovery
```

Track progress against this checklist. Move on only when each step verifies.

> **Log every stage.** A local integration has many silent failure points —
> spawn, health, backend install, model download, first inference. Without a
> log line at each transition, "nothing happened" is indistinguishable from
> "broke at stage 3." Emit one clear line per stage as you build (see
> [Step 4](#step-4-add-a-lemond-launcher)); the most common dead-end in this
> integration — a blank result with no error — is invisible without them.

---

## Step 1: Survey the app

Find every place the app currently calls a cloud AI API. Search the repo for:

- `openai`, `OpenAI(`, `chat.completions`, `responses.create`
- `anthropic`, `Anthropic(`, `messages.create`
- `api.openai.com`, `api.anthropic.com`, `localhost:11434` (Ollama)
- `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`

Record three things before continuing:

1. **Client library and language** (e.g., `openai-python`, `openai-node`,
   `@anthropic-ai/sdk`, `go-openai`, raw `fetch`).
2. **Modalities used:** text chat, tool calling, embeddings, image gen,
   transcription, TTS. This drives the model + backend choice in Step 2.
3. **One single place** where the base URL and API key are constructed. If
   there isn't one, refactor to one before going further. Local-mode toggling
   must flip exactly one config object.
4. **Any API-key gating** that blocks the app before a key is entered
   (onboarding walls, validators that reject empty keys, startup checks that
   disable AI until a key exists). Note each one — Step 5 bypasses them in
   local mode.

## Step 2: Pick a model + backend profile

Choose **one** default profile based on the app's primary modality. Do not
ship a buffet. Ship one good default and document how the user can override
it.

| App's primary need | Default model | Recipe | Why |
|---|---|---|---|
| General chat / assistant | `Qwen3-4B-GGUF` | `llamacpp` | Small, fast, good tool calling, fits 8GB systems |
| Coding assistant | `Qwen2.5-Coder-7B-Instruct-GGUF` | `llamacpp` | Strong code, runs on iGPU |
| Vision / multimodal chat | `Gemma-4-E2B-it-GGUF` | `llamacpp` | Small multimodal default |
| NPU-first on Ryzen AI | `Llama-3.2-3B-Instruct-Hybrid` | `ryzenai-llm` | XDNA2 NPU on Windows |
| Speech-to-text (Windows) | `Whisper-Large-v3-Turbo` | `whispercpp` | One model; probe picks NPU → iGPU/dGPU → CPU automatically |
| Speech-to-text (Linux NPU) | `whisper-v3-turbo-FLM` | `flm` | Linux NPU path; falls back to `whispercpp` iGPU/CPU off-NPU |
| Text-to-speech | `kokoro-v1` | `kokoro` | CPU-only, low latency |
| Image generation | `SDXL-Turbo` | `sd-cpp` | Single-step generation |

For the LLM backend, default to `llamacpp` and let `lemond` pick
`rocm` → `vulkan` → `cpu` automatically by leaving `llamacpp_backend`
unset. Override only if the app has hard hardware requirements.

**Scope: this skill selects a backend once at integration time on the
developer's machine.** Runtime fallback based on the end user's hardware is
out of scope. Bundle `vulkan` as the universal fallback so the app works on
any machine. If the dev machine has an NPU and the chosen recipe supports it,
the skill will use the NPU backend — otherwise it falls back to `vulkan`.

> **Note:** having an NPU does not mean every recipe supports NPU. Confirm
> the recipe/backend pair is `installed` or `installable` via
> `GET /api/v1/system-info` before committing to it. See
> [reference.md](reference.md#hardware-probing-with-v1system-info) for
> per-recipe decision rules.

For more options and tradeoffs, see [reference.md](reference.md).

## Step 3: Place Embeddable Lemonade in the app's tree and install backends

**Get the embeddable artifact** from the latest Lemonade release:

```
https://github.com/lemonade-sdk/lemonade/releases/latest
```

Download the file matching your target OS:

- Windows: `lemonade-embeddable-{VERSION}-windows-x64.zip`
- Linux:   `lemonade-embeddable-{VERSION}-ubuntu-x64.tar.gz`

> **Don't hand-build the download URL from the tag.** The git tag carries a
> leading `v` (e.g. `v10.8.0`) but the asset filename strips it
> (`lemonade-embeddable-10.8.0-...`), so using the tag verbatim 404s. Ask the
> GitHub API for the asset by its stable name pattern and use the URL it
> returns, as below — this stays correct across version and naming changes.

**First, create the target directory** — it does not exist in a fresh repo:

```powershell
# Windows
New-Item -ItemType Directory -Force vendor\lemonade
```

```bash
# Linux
mkdir -p vendor/lemonade
```

Then download and unpack on Windows (PowerShell):

```powershell
$rel = Invoke-RestMethod https://api.github.com/repos/lemonade-sdk/lemonade/releases/latest
$asset = $rel.assets | Where-Object { $_.name -like "lemonade-embeddable-*-windows-x64.zip" } | Select-Object -First 1
Invoke-WebRequest $asset.browser_download_url -OutFile lemond.zip
Expand-Archive lemond.zip -DestinationPath "$env:TEMP\lemond-unpack"
$folder = $asset.name -replace '\.zip$',''   # unpacked dir = asset name without .zip
Copy-Item -Recurse "$env:TEMP\lemond-unpack\$folder\*" vendor\lemonade\
# Sanity check: resources/ must be nested under vendor\lemonade\ (not flattened)
if (-not (Test-Path vendor\lemonade\resources\*.json)) { throw "resources/ missing — re-extract and copy again" }
```

On Linux (bash):

```bash
URL=$(curl -s https://api.github.com/repos/lemonade-sdk/lemonade/releases/latest \
  | grep browser_download_url | grep ubuntu-x64.tar.gz | cut -d'"' -f4)
curl -L "$URL" | tar -xz --strip-components=1 -C vendor/lemonade
```

> **Copy the full package, not just the binary.** The archive contains
> `lemond[.exe]`, `lemonade[.exe]`, `LICENSE`, and `resources/`. The
> `resources/` directory is required — without it lemond starts and passes the
> health check but fails on every model and backend request. Copying only the
> binary produces a server that looks healthy but cannot function.

> **`lemond` vs `lemonade` CLI:** `lemond` is the embedded server binary that
> ships with the app. The `lemonade` CLI is a separate packaging tool used
> only during development/build time to install backends. The same embeddable
> archive unpacked above already contains a matching `lemonade[.exe]` next to
> `lemond[.exe]`, so its version aligns with the bundled `lemond`. Do **not**
> `pip install lemonade-sdk` to get it: the PyPI package is a separate, older
> release line whose ports, model names, and install API do not match the
> `lemond` bundled here, and mixing the two is a known source of silent
> version mismatches. Keep the `lemonade` CLI, `lemond`, and the backends all
> from the one release downloaded in this step so their versions stay aligned.

The expected layout **after setup** (first run + backend install). A freshly
unzipped package contains only `lemond[.exe]`, `lemonade[.exe]`, `LICENSE`, and
`resources/` — the items below are created later, as their comments note:

```
vendor/lemonade/
  lemond[.exe]                     # the only binary the app ships
  LICENSE
  config.json                      # generated on first run; commit a seed copy
  resources/
    server_models.json             # do not edit; use GET /api/v1/models at runtime
    backend_versions.json
  bin/                             # backends bundled at packaging time
    llamacpp/vulkan/llama-server[.exe]
  models/                          # pre-bundled model weights (optional)
    models--unsloth--Qwen3-4B-GGUF/
```

> **`server_models.json`:** Do not edit or rely on this file. It can be stale.
> The only authoritative model list is `GET /api/v1/models` on a running
> `lemond` instance with the backend already installed.

**Bundle decisions: pick deliberately**

- **Backends:** Bundle `llamacpp:vulkan` at packaging time (works on every
  GPU). Install `llamacpp:rocm` at first run on supported AMD systems via
  `POST /api/v1/install` after probing `GET /api/v1/system-info`. Never ship
  every backend, or the artifact balloons.
- **Models:** Either bundle the default model under `models/` (offline
  install, larger installer) **or** pull on first run with
  `POST /api/v1/pull` (smaller installer, needs network). Pick one and
  document it.
- **`models_dir`:** Set to `./models` in `config.json` to keep weights
  private to the app. Leave as `auto` only if the user explicitly wants to
  share weights with other apps.

**Backend install timing — two distinct paths:**

> **Packaging time** (developer machine, before bundling). Use the lemonade
> CLI that shipped inside `vendor/lemonade/` so it matches the bundled
> `lemond` version (prefix with `./` or the full path):
> ```
> vendor/lemonade/lemonade backends install llamacpp:vulkan
> vendor/lemonade/lemonade backends install flm:npu    # Windows NPU path only
> ```
> This bakes the backend binaries into `vendor/lemonade/bin/` before the app
> ships. `lemond` does not need to be running. Use a modern `lemonade` CLI
> whose version matches the bundled `lemond` (the copy in the archive you
> unpacked works); do not `pip install lemonade-sdk` for it.
>
> **First-run / runtime** (user's machine, after `lemond` is running):
> ```http
> POST /api/v1/install
> {"recipe": "llamacpp", "backend": "rocm"}
> ```
> Use this for hardware-specific backends (e.g. `llamacpp:rocm`) that cannot
> be bundled universally. `lemond` must already be running (Step 4 complete).

## Step 4: Add a `lemond` launcher

Write the launcher as a new module named **`lemond_launcher.py`** (or
`lemond_launcher.<ext>` for the app's language). It is a thin process
supervisor. Its only jobs:

1. Generate a fresh random API key: `key = secrets.token_urlsafe(32)`
2. Pick a free localhost port: bind a `socket` to port 0, read back the assigned port, close it.
3. Spawn lemond as a `subprocess`: `subprocess.Popen([LEMOND_BIN, LEMOND_DIR, "--port", str(port)], env={**os.environ, "LEMONADE_API_KEY": key})`
4. Poll `GET /api/v1/health` with `Authorization: Bearer {key}` in a loop until HTTP 200 — this is the only correct readiness check.
5. Expose the chosen `port` and `key` to the rest of the app.

> **Log one line per lifecycle stage.** Build the logging in from the start —
> not as an afterthought when something breaks. Each silent transition needs a
> visible marker so a failure points at the exact stage. Aim for:
>
> ```
> [lemond] Starting on port <port>
> [lemond] Healthy on port <port>
> [lemond] <recipe>:<backend> installed        (or: already installed / install failed)
> [lemond] Pulling model <name>...             then: Model <name> ready  (or: pull returned <status>)
> [local]  <modality> result: <value>          (first inference output — empty string here = unpulled model)
> ```
>
> Logging the **first inference result verbatim** is what turns the
> silent-empty failure (Step 6) from a multi-hour mystery into a one-line
> diagnosis. Route these through the app's normal logging so they can be quieted
> for release.

> **Dev-mode file watchers:** If the app runs with a file watcher (Tauri,
> Electron, Next.js, Vite, etc.) that watches the source tree, ensure
> `vendor/lemonade/` is excluded from the watched paths. Lemond writes config
> and cache files at runtime; a watcher that picks these up will restart the
> app, kill the lemond subprocess, and spawn a new one on a new port —
> silently breaking any in-flight transcription. Add `vendor/` (or the
> equivalent) to the watcher's ignore list before testing.

**Use the reference implementation from [reference.md § Reference launchers](reference.md#reference-launchers) directly** — copy it verbatim and adapt only the `LEMOND_DIR` path. Do not write a launcher from scratch. The reference Python launcher uses `secrets` (for the API key), `socket` (for the free-port probe), and `subprocess` (to spawn lemond); the Node.js launcher uses the equivalent stdlib modules. Both handle port-race retries and health polling correctly.

Readiness is always determined by polling the exact endpoint
`GET http://127.0.0.1:<port>/api/v1/health` and checking for HTTP 200 — never
by reading `lemond`'s stdout or stderr. Any health-check helper you write must
hit that `/api/v1/health` path.

## Step 5: Re-point the existing client at `lemond`

Make **three** changes to the app's existing client construction — all three
are required, not optional:

1. Set `base_url` to `http://127.0.0.1:{port}/api/v1`
2. Set `api_key` to the launcher key
3. **Set the HTTP timeout to 120 seconds** — this is mandatory, not optional

The 120-second timeout is not a tuning suggestion. The default on most HTTP
clients is 30s, which is shorter than lemond's first-run model load time on
real hardware. Without it the request silently times out and the UI shows
nothing, which is indistinguishable from a broken integration.

**Python (openai) — the exact change to make:**

```python
import httpx
from openai import OpenAI

proc, key, port = start_lemond()
client = OpenAI(
    base_url=f"http://127.0.0.1:{port}/api/v1",
    api_key=key,
    http_client=httpx.Client(timeout=120),  # required: 120s for first-run model load
)
```

For other clients:

| Existing client | New `base_url` | New auth | Timeout |
|---|---|---|---|
| `openai-python` | `http://127.0.0.1:{port}/api/v1` | `api_key=key` | `httpx.Client(timeout=120)` |
| `openai-node` | `http://127.0.0.1:{port}/api/v1` | `apiKey: key` | `timeout: 120000` |
| `@anthropic-ai/sdk` | `http://127.0.0.1:{port}/api/v1` | `apiKey: key` | `timeout: 120000` |
| Raw `fetch` / `requests` | same | `Authorization: Bearer {key}` | set per-request |
| Ollama-compatible code | `http://127.0.0.1:{port}/api/v0` | pass key anyway | 120s |

The model identifier on requests stays a Lemonade model name (e.g.
`Qwen3-4B-GGUF`), not the cloud name.

**Local mode needs no cloud API key — at all.** This is a defining property of
local mode, not an edge case: there is no cloud service to authenticate to, so
nothing should ever ask the user for a key. Any onboarding wall, validator, or
startup check that demands one must not block local-mode users. Concretely:

- Skip or auto-satisfy the key-entry screen in local mode.
- Treat local mode as already-authorized in every validation path — an
  empty-key check must short-circuit to "valid" when the active mode is local,
  never throw "API key not configured".
- Re-enable the gate **only** for cloud mode.

The `lemond` key from Step 4 is generated internally by the launcher and used
only for the local loopback connection, so the user never sees or enters one;
any UI placeholder (e.g. `"local"`) is fine. Flipping into local mode should
never strand the user on a key-entry wall.

## Step 6: Health, backend, then pull the model — *before* first inference

`GET /api/v1/health` returning 200 means the **server** is up. It does **not**
mean inference will work. Before the first real request succeeds, three more
things must be true: the backend for your modality is installed, the model's
weights are **downloaded to disk**, and (on the first call) the model is loaded
into memory. Treating health=200 as "ready" is the single biggest cause of a
broken-looking integration.

**Do not call `POST /api/v1/load` at startup.** Lemond lazy-loads the model
into memory on the first inference request and handles that step on its own.
Pre-loading is unreliable across lemond versions (the `/load` request body
shape has changed between releases) and a malformed call can crash or
destabilise the server before the user takes any action. Loading is the one
step you let lemond do lazily — pulling is not.

### Pull the model so it exists on disk

Lazy-load only loads weights that are **already downloaded**. If the model was
never pulled, the first inference does not error — lemond returns an empty /
blank result with HTTP 200. So after health passes and the backend is
installed, proactively pull the model:

```http
POST /api/v1/pull
{"model": "Whisper-Large-v3-Turbo"}
```

This is **idempotent** — a no-op if the weights are already present, a download
if they are not. Run it once during setup (after backend install, before the
first user-triggered inference) and log the result.

- **Default model** (the one you chose in Step 2): pull it by name as above.
- **Custom / user-overridden model:** do not assume it exists. Confirm it is a
  real Lemonade model first via `GET /api/v1/models` (the **only** trusted
  catalog — see [reference.md](reference.md)), then pull it the same way. A
  model appearing in the catalog is **not** proof its weights are downloaded;
  a successful pull is.

> **Silent-empty is almost always an unpulled model.** If inference returns an
> empty string / blank output with no HTTP error, the model was not downloaded.
> Check your pull step before debugging anything else — this is the failure mode
> that wastes the most time. Log the pull result and the first inference result
> (see Step 4) so this is diagnosable from the console, not by guesswork.

### Surface the *whole* setup, not just model load

First-run cold start is more than a model load. The full sequence is:

```
server spawn  →  health 200  →  backend install  →  model download  →  model load  →  first result
```

On a fresh machine, backend install and model download can each take from tens
of seconds to several **minutes** (multi-GB weights over the network). Model
load alone is 10–30s. An app that shows nothing during this will look frozen.

Minimum: show a loading indicator or status message ("Setting up local AI…")
from the moment setup begins until the first response arrives — covering the
*entire* sequence above, not just the final load. The simplest implementation
is a flag set when setup/first-request starts and cleared when the first
response arrives. Once the model is pulled and loaded once, subsequent runs are
fast; the long wait is first-run only.

## Step 7: Lifecycle and recovery

These are the only failure modes worth handling. Do not over-engineer.

| Symptom | Cause | Recovery |
|---|---|---|
| **Inference returns empty / blank with HTTP 200, no error** | Model never pulled: backend is installed but weights are absent, so lazy-load has nothing to load | `POST /api/v1/pull` with `{"model":"..."}`, wait for success, retry. Log the pulled result and the first inference result. This is the most common silent failure — see [Step 6](#step-6-health-backend-then-pull-the-model--before-first-inference) |
| `POST /api/v1/load` returns 404 / model not found | Model not pulled yet (same root cause as the empty-result row above) | `POST /api/v1/pull` with `{"model": "..."}` then retry `/api/v1/load` |
| `POST /api/v1/load` returns 500 with backend error | Backend not installed for this hardware | `GET /api/v1/system-info`, pick a supported backend, `POST /api/v1/install` with `{"recipe": "...", "backend": "..."}`, retry |
| Subprocess exits immediately | Port race: another process grabbed the port between `freePort()` and lemond binding | The reference launcher retries with a fresh port automatically (3 attempts) |
| `/api/v1/health` never returns 200 | First-run backend extraction is slow on cold disk | Extend timeout to 90s on first launch, 30s after |
| HTTP 401 on every request | Forgot the `Authorization: Bearer` header | Audit the client config because Lemonade rejects unauth'd calls when `LEMONADE_API_KEY` is set |

**Shutdown:** On app exit, `proc.terminate()` (Unix) or
`proc.kill()` (Windows). `lemond` flushes config and exits cleanly within a
couple of seconds. Always wait on the process; never orphan it.

**Do not** parse `lemond` stdout to detect readiness; use the HTTP
`/api/v1/health` probe. Stdout format is not a stable contract.

---

## Verification checklist

The integration is done when **all** of these are true:

- [ ] `vendor/lemonade/` contains the full package: `lemond[.exe]`,
      `lemonade[.exe]`, `LICENSE`, and `resources/` — not just the binary.
- [ ] `lemond` starts as a subprocess with a fresh API key per launch.
- [ ] `GET /api/v1/health` returns 200 within the timeout.
- [ ] The default model is pulled (or bundled) before the first inference; a
      custom/overridden model is confirmed via `GET /api/v1/models` and then
      pulled. A blank result with no error means this step was skipped.
- [ ] Each lifecycle stage logs a clear line (spawn, health, backend install,
      model pull, first result) so a failure is diagnosable from the console.
- [ ] The existing client's chat / image / speech call returns a valid
      response with the base URL and key swapped, with no other code changed.
- [ ] First-run latency is surfaced: the interface shows a loading state from the
      moment the first inference request is sent until the response arrives.
- [ ] The HTTP client timeout is set to 120 seconds.
- [ ] In local mode the app requires **no** cloud API key: no onboarding wall,
      validator, or startup check blocks the user, and no code path throws
      "API key not configured" when the active mode is local.
- [ ] If the app uses a dev-mode file watcher, `vendor/lemonade/` is excluded
      from the watched paths so runtime writes by lemond do not trigger restarts.
- [ ] Killing the parent process leaves no `lemond` subprocess behind.
- [ ] On a fresh machine without the optimal backend, the app still works
      via the Vulkan fallback bundled in `bin/`.

If any box is unchecked, do not declare the task complete.

---

## Reference

For detailed model catalog, backend selection matrix, full endpoint reference,
config keys, and per-model `recipe_options.json` tuning, see
[reference.md](reference.md).

Referenced files: 5

local-ai-use17.8 KB

View saved version →

---
name: local-ai-use
description: >-
  Makes this agent generate images, transcribe audio, and synthesize speech on
  the user's own machine through a local Lemonade Server instead of a paid cloud
  API. Use it above all to change that routing persistently, from now on — keep
  generating pictures locally while chat stays on the cloud; set this workspace
  up to make images on my own machine — even when the user asks for no image or
  file in the same breath. Also use it for a single request the user wants done
  locally, offline, on-device, or kept private: transcribe this recording, make
  this picture, read this text aloud. Applies in Claude, Cursor, Codex, or any
  agent harness. Use when the user wants to cut cost or tokens on image, audio,
  or voice API calls, or to drop DALL-E, hosted Whisper, ElevenLabs, or other
  paid multimodal APIs; or mentions Lemonade Server, OmniRouter, SD-Turbo,
  kokoro, Ryzen AI, or NPU/iGPU/dGPU inference. Changes no application source
  code; do not use it if the user is adding local AI to an app they ship.
---

# Local AI Use (route image, TTS, STT through Lemonade)

This is a **meta-skill**. You run it once. After that, every later request that
needs image generation, text-to-speech, or speech-to-text uses the local
[Lemonade Server](https://lemonade-server.ai) instead of a cloud API. The
agent's own LLM keeps handling text; only the expensive multimodal calls move
on-device.

The skill does three things:

1. **Makes sure local Lemonade is installed and running.** If no modern
   `lemonade` CLI is found, the setup script installs the latest version of
   Lemonade on the user's behalf. Modern Lemonade has no `serve` command — the
   Lemonade service (the `lemond` daemon) auto-starts on install and is managed
   by the OS — so the setup script waits for the service and, if it stays down,
   prints the exact OS-specific command to start it (e.g. `sudo systemctl start
   lemond` on Linux).
2. **Verifies that local Lemonade is reachable.**
3. **Drops a `Local AI Use` block into the workspace `AGENTS.md`** so the agent
   reads the routing rule on every later turn, in Cursor, Claude Code, Codex,
   Gemini CLI, and any other agent that respects `AGENTS.md`.

> **Requires modern Lemonade (v10.1.0 or newer).** Modern Lemonade unified
> everything under one `lemonade` CLI (`lemonade status`, `lemonade pull`, ...)
> driving an always-on `lemond` service. `lemonade` is the only valid CLI, and
> this skill installs Lemonade only from the [official install
> paths](https://lemonade-server.ai/docs/guide/install/) listed in Step 1a. If
> an older `lemonade` is already on the `PATH` it will shadow the modern CLI;
> uninstall it first (see the removal commands in Step 1a) before running this
> skill.

Models are **not** downloaded during setup. Each default model is pulled
lazily, on first use, by the routing rule (e.g. the first image request pulls
the image model). This keeps setup fast and avoids gigabytes of downloads the
user may never need.

## When to use this skill

Use this skill when **all** of the following are true:

- The user wants local Lemonade. If it is not yet installed, the setup script
  installs the latest version for them automatically.
- The user accepts the default Lemonade endpoint `http://localhost:13305`.
- The user wants the change to be **persistent** across future turns and
  agent restarts (the rule is written to disk).

If the user is instead **embedding** Lemonade as a private subprocess inside
an app installer, do not use this skill; use `local-ai-app-integration`
instead.

## Prerequisites

- **OS:** Windows 11 x64, Ubuntu/Debian x64, or macOS (beta).
- **Lemonade:** the setup script installs the latest version if missing, using
  `winget` on Windows, the `ppa:lemonade-team/stable` PPA on Ubuntu/Debian, and
  the Homebrew cask on macOS (see Step 1a for the fallbacks). The `lemond`
  service auto-starts after install; the script waits for it rather than
  launching it. On Linux the install needs `sudo`. Pass `--no-install` if the
  user wants to install it themselves instead.
- **Disk:** ~8 GB free for the three default models (SD-Turbo + Whisper-Tiny
  + kokoro-v1), plus ~0.1 GB for the installer itself. The first image request
  also triggers a ~5 GB pull for `SD-Turbo` if it is not already cached; on
  metered or slow links, consider pulling models eagerly after setup (see
  `lemonade pull` in `reference.md`).
- **Network:** required for the install download and the first `lemonade pull`
  of each model. After that, every modality runs offline.
- **Version:** requires v10.1.0 or newer (the unified `lemonade` CLI this
  skill targets). Model IDs and `system-info` fields can change between
  releases; confirm against `lemonade status` and `GET /api/v1/models` on the
  version actually installed rather than assuming this document is current.

## The opinionated path

Run this checklist top to bottom. Track progress against it; do not move on
until each step verifies.

```
[ ] 1. Ensure Lemonade Server is installed and running (auto-install if missing)
[ ] 2. Install the routing rule into the workspace AGENTS.md
```

On a managed or shared machine where the agent must not run
`sudo apt-get install`, pass `--no-install` to the setup script and confirm
Lemonade is already installed before continuing.

The single command that does both steps in one shot is:

```bash
python scripts/setup_local_ai.py
```

**Always run this script first — even if Lemonade is already installed and the
server is already running, and even before generating a single image.** Writing
the routing rule into `AGENTS.md` is what makes this skill complete; skipping it
because "Lemonade is already up" leaves the workspace unconfigured for future
turns. The script is safe to run in that case: it detects the running service,
skips the install, and just writes the rule.

It auto-installs the latest version of Lemonade if no modern `lemonade` CLI
is found, waits for the auto-started `lemond` service, then writes the rule.
The script is idempotent: re-running it on a fully configured workspace is a
no-op apart from a healthcheck. Read the sections below for what to do when
each step fails.

---

## Step 1: ensure Lemonade Server is installed and running

`scripts/setup_local_ai.py` handles this end to end, but here is what it does
so you can do it by hand or debug it. Pass `--no-install` when Lemonade is
already managed elsewhere and the agent must not attempt a package install.

**1a. Is a modern `lemonade` CLI installed?** Run `lemonade status`. The check
is by *capability*, not by name: modern Lemonade prints `Server is running...`
or `Server is not running`. If instead you get an "invalid choice" / usage
error, the `lemonade` on `PATH` is an old build that predates the unified CLI
(v10.1.0) — do **not** use it. Remove it, then re-run this skill:

| OS | Uninstall the old build with |
|---|---|
| Windows | `winget uninstall -e --id AMD.LemonadeServer`, or Settings > Apps > Installed apps > Lemonade Server > Uninstall |
| Ubuntu/Debian | `sudo apt remove lemonade-server` |
| macOS | `brew uninstall --cask lemonade-server`, or delete the installed `Lemonade.app` and its `.pkg` receipt |

Never try to drive or auto-remove it for the user.

If no `lemonade` is found at all, install the latest version on the user's
behalf. Use the package manager first; the download is the fallback when the
package manager is absent. Full matrix, including Arch, Fedora, Debian, Snap,
and Docker, is in the [install docs](https://lemonade-server.ai/docs/guide/install/).

| OS | Install | Fallback |
|---|---|---|
| Windows | `winget install -e --id AMD.LemonadeServer` | Download [`lemonade.msi`](https://github.com/lemonade-sdk/lemonade/releases/latest/download/lemonade.msi) and run `msiexec /i lemonade.msi /qn` (silent, per-user, no elevation). |
| Ubuntu | `sudo add-apt-repository -y ppa:lemonade-team/stable && sudo apt-get update && sudo apt-get install -y lemonade-server` | `sudo snap install lemonade-server` |
| macOS | `brew install --cask lemonade-server` | Download `Lemonade-<ver>-Darwin.pkg` from the [latest release](https://github.com/lemonade-sdk/lemonade/releases/latest) and run `sudo installer -pkg Lemonade-<ver>-Darwin.pkg -target /`. |

The Ubuntu apt package is named `lemonade-server`, but the CLI it installs is
`lemonade`. The browser UI is served at `http://localhost:13305` with no extra
package; add `sudo apt install lemonade-desktop` only if the user wants the
desktop frontend.

After a Windows install the CLI lands in `%LOCALAPPDATA%\lemonade_server` and
is added to the *user* PATH (new shells only); the setup script probes that
directory so it works in the same run.

**1b. Is the service running?** Check `lemonade status --json`. The `lemond`
service auto-starts on install — there is **no** `lemonade serve` in modern
Lemonade.

| `lemonade status` says | Action |
|---|---|
| `Server is running on port 13305` | Continue to Step 2. |
| `Server is not running` | Wait a few seconds for the auto-started service (the script polls `/api/v1/health`). If it stays down, start it via the OS service manager: `sudo systemctl start lemond` (Linux system install) or `systemctl --user start lemond` (per-user install); `launchctl load /Library/LaunchDaemons/com.lemonade.server.plist` (macOS); the Lemonade tray app or `Start-Service lemond` (Windows). |

Only if the automatic install genuinely fails (no `apt-get`, no `sudo`,
download blocked) should you stop and point the user at
<https://lemonade-server.ai/docs/guide/install/>.

The rest of this skill assumes the endpoint is `http://localhost:13305/api/v1`
and no API key is required (the system-wide server defaults to no auth on
loopback). If the user has set `LEMONADE_API_KEY`, the routing rule template
in `templates/local-ai-rule.md` shows where to add the `Authorization` header.

**1c. Are the backends ready per modality?** Backend health is **per
modality**. A working chat or image request does not prove transcription will
work: `auto` picks a different backend per modality and can silently fall back
for one while having no alternative for another. Before declaring setup
complete, check the actual per-modality state:

```bash
lemonade backends --all
```

Any variant the workspace's routing depends on should read `installed`. If the
only installed variant for a modality is `rocm`, install the Vulkan variant as
well so `auto` has somewhere to fall back to (for example,
`lemonade backends install whispercpp:vulkan`).

### Default modality models (pulled on first use, not during setup)

Setup does **not** download these. The installed rule pulls each one the first
time that modality is requested. They are the smallest models Lemonade offers
per modality, sized to keep token-and-cost savings real on commodity hardware:

| Modality | Model | Size | Why this default |
|---|---|---|---|
| Image generation | `SD-Turbo` | ~5 GB | Single-step generation, runs on CPU and AMD iGPU/dGPU |
| Text-to-speech | `kokoro-v1` | ~0.3 GB | Only TTS model Lemonade currently supports; CPU-only, low latency |
| Speech-to-text | `Whisper-Tiny` | ~0.1 GB | Smallest Whisper; fast on CPU. Upgrade to `Whisper-Large-v3-Turbo` if accuracy matters more than latency. |

To write a different model ID into the rule, pass it to the setup script. For
example, to make future image requests use SDXL:

```bash
python scripts/setup_local_ai.py --image-model SDXL-Turbo
```

That model ID is written into the installed `AGENTS.md` rule and pulled on its
first use. The same pattern works for `--tts-model` and `--stt-model`. For
larger / higher-quality alternatives (`SDXL-Turbo`, `Flux-2-Klein-4B`,
`Whisper-Large-v3-Turbo`), see the
[model picker in reference.md](reference.md#model-picker).

## Step 2: install the routing rule into AGENTS.md

The rule is a Markdown block stored in [`templates/local-ai-rule.md`](templates/local-ai-rule.md).
Append it to the workspace's `AGENTS.md` (create the file if missing). Both
Cursor and Claude Code load `AGENTS.md` automatically on every turn, so the
agent will see the rule on its next message without any further setup.

`scripts/setup_local_ai.py` does this for you. It bakes the selected endpoint
and model IDs into the rule, surrounded by stable markers so re-running the
script replaces the block in place rather than appending a second copy. The
markers look like:

```
<!-- BEGIN amd-skills:local-ai-use -->
...rule...
<!-- END amd-skills:local-ai-use -->
```

If you write the file by hand, keep those exact markers. The script relies
on them for idempotent updates.

If the user's agent only respects a different convention, mirror the same
block to:

- `CLAUDE.md` (Claude Code, project-scoped) or `~/.claude/CLAUDE.md` (global)
- `.cursor/rules/local-ai-use.mdc` (Cursor user/project rules)
- `GEMINI.md` (Gemini CLI)

The rule's content is identical; only the file location changes.

---

## What changes after this skill runs

From the next turn onward, the agent reads the rule in `AGENTS.md` on every
message. The rule explicitly tells the agent:

- **For image generation:** call `POST /api/v1/images/generations` on the
  local server. Do **not** call any cloud image API and do **not** use the
  built-in `GenerateImage` tool (that path bills tokens to the cloud
  provider).
- **For text-to-speech:** call `POST /api/v1/audio/speech`. Do **not** call
  cloud TTS providers (OpenAI TTS, ElevenLabs, etc.).
- **For speech-to-text:** call `POST /api/v1/audio/transcriptions`. Do
  **not** call cloud transcription providers.
- **Fallback:** only fall back to a cloud API after one local attempt has
  failed *and* the user has been told the local call failed. Never silently;
  the whole point of this skill is to keep cost predictable. For
  speech-to-text, the disclosure must also say the transcript came from a
  different engine, since mixed-engine transcripts should not be compared or
  deduplicated.

The agent's own text reasoning continues to use whatever LLM Cursor / Claude
Code / Codex is configured with. This skill does not redirect chat tokens;
it only redirects the multimodal calls that would otherwise leave the
machine.

## Troubleshooting cheatsheet

| Symptom | Cause | Recovery |
|---|---|---|
| `lemonade: command not found` | CLI not installed | Re-run `python scripts/setup_local_ai.py` (auto-installs the latest version). If it just installed on Windows, open a new shell so the user PATH refreshes, or the script will find it under `%LOCALAPPDATA%\lemonade_server`. |
| `status` gives an "invalid choice" / usage error | An old `lemonade` (pre-v10.1.0) is shadowing the modern CLI | Uninstall it (see the Step 1a table: `winget uninstall -e --id AMD.LemonadeServer` / `sudo apt remove lemonade-server` / `brew uninstall --cask lemonade-server`), then re-run the setup script. |
| `Server is not running` | `lemond` service stopped | Start it via the OS service manager — `sudo systemctl start lemond` / `systemctl --user start lemond` (Linux), `launchctl load /Library/LaunchDaemons/com.lemonade.server.plist` (macOS), or the tray app / `Start-Service lemond` (Windows). There is no `lemonade serve`. |
| `POST /v1/images/generations` returns 404 model not found | Image model not downloaded | `lemonade pull SD-Turbo` and retry. |
| `lemonade pull` keeps printing `Progress: NN%` but never finishes | Download target is a bad path (out of space, no write permission, quota, read-only mount). The write error may surface only in the server log while the console keeps showing progress | Check the target and free space first: `GET /api/v1/system-info` reports `models_dir` and `model_storage.free_bytes`. If a pull stalls, read the recent lines of the server log (typically `lemonade-server.log` in the OS temp dir) for the real error (e.g. a download/write failure like `CURL code 23`, or an out-of-space message), then point the download at a writable disk with room. |
| Image generation is slow on CPU (~4–5 min) | sd-cpp on CPU backend | Install the GPU backend on supported AMD hardware: `lemonade backends install sd-cpp:rocm`. |
| Still slow after installing the GPU backend | The backend is installed but not actually engaged; the runtime fell back to CPU silently | An `installed` state in `system-info` and a successful `rocminfo` both still permit a silent CPU fallback. Check real GPU utilisation (`gpu_busy_percent`) during a request, and confirm the host's GPU driver stack rather than re-installing the backend. |
| `POST /v1/audio/transcriptions` returns 400 unsupported format | Input is not 16 kHz mono WAV | Re-encode with `ffmpeg -i in.* -ar 16000 -ac 1 out.wav`. |
| `POST /v1/audio/speech` returns 404 | TTS model not downloaded | `lemonade pull kokoro-v1`. |
| 401 Unauthorized on every request | User has set `LEMONADE_API_KEY` | Add `Authorization: Bearer $LEMONADE_API_KEY` to every request and to the rule block. |

## Verification checklist

Mark this skill complete only when **all** of the following are true:

- [ ] `lemonade status --json` reports the server running on port 13305.
- [ ] The workspace `AGENTS.md` contains the
      `amd-skills:local-ai-use` block. This is required even when Lemonade was
      already installed and running — generating an image alone does not
      complete the skill.
- [ ] On a follow-up turn, asking the agent to "generate an image of X"
      causes it to POST to `http://localhost:13305/api/v1/images/generations`
      (pulling the model on first use) rather than calling a cloud tool.
- [ ] `lemonade backends --all` shows `installed` for every backend variant
      this workspace's routing depends on (see Step 1c). Do not treat a working
      image or chat path as proof that transcription will work.

If any box is unchecked, the user is still paying cloud cost for at least
one modality, or a routed modality may fail silently on first use.

---

## Reference

For the full model picker, alternate-quality options, the complete endpoint
reference, the API-key flow, and the OmniRouter tool definitions you can
hand to an agent's tool-calling loop, see [reference.md](reference.md).

Referenced files: 5

serving-llms-on-instinct15.5 KB

View saved version →

---
name: serving-llms-on-instinct
description: >-
  Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill
  whenever the user wants to run, serve, deploy, start, host, or launch a
  language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X.
  Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM,
  or asks how to get a model running on AMD data center hardware. Use when the
  user asks "run Qwen3", "serve DeepSeek", "start a vLLM endpoint", "get a
  model running on my AMD machine", or any similar phrasing. Handles the full
  flow: GPU detection, environment validation, vLLM configuration, launch, and
  health verification. Do not use for NVIDIA GPUs, consumer AMD GPUs (RX
  series, Radeon), Ryzen AI, NPU, MI250X, or MI100.
allowed-tools: Bash, Read
---

# Serving LLMs on AMD Instinct

Get a vLLM endpoint running on AMD Instinct GPU hardware.

## Prerequisites

- ROCm driver and `amd-smi` installed on the GPU host
- Docker running and accessible (check with `docker ps`)
- `/dev/kfd` and `/dev/dri` present on the GPU host
- HuggingFace token in `HF_TOKEN` env var (required for gated models; not
  required for Qwen3 or Gemma). For gated models (Llama 3.2, Gemma, etc.),
  the HF token must belong to an account that has accepted the model's license
  at `huggingface.co/<model_id>`. A valid token without license acceptance will
  fail with an opaque "Engine core initialization failed" error.
- For remote GPU: SSH key access configured (`ssh <user>@<host>` must work
  without a password prompt). If only password access is available, set up
  keys first: `ssh-copy-id <user>@<host>`

## Data files

Read these files directly to get model and GPU configuration:

- **`data/recipes_cache.json`** -- model configs synced from
  [vllm-project/recipes](https://github.com/vllm-project/recipes). Each entry
  under `models.<HF_ID>.recipe` contains the full recipe with `model.base_args`,
  `model.base_env`, `features.tool_calling.args`, `features.reasoning.args`,
  `hardware_overrides.amd.extra_args`, `hardware_overrides.amd.extra_env`.
  The top-level `docker_image` field has the latest resolved vLLM ROCm image.

- **`data/gpu_overrides.json`** -- GPU-specific configuration. Contains
  `docker_flags` (mandatory for all AMD Instinct), `gpu_configs` keyed by
  gfx_version with `env_defaults` and `workarounds`, and `legacy_models` for
  models not yet in vLLM recipes.

- **`data/blacklist.json`** -- models in vLLM recipes that cannot be served
  as LLM endpoints. Includes diffusion/image/audio generation models, embedding
  models, rerankers, ASR models needing audio pipelines, and models requiring
  unreleased vLLM nightly builds. Check this before attempting to serve a model.
  If the user requests a blacklisted model, explain why it won't work and
  suggest an alternative.

If the user doesn't specify a model, default to **Qwen/Qwen3.5-9B**: dense
multimodal with MTP, Apache 2.0 license (no HF token needed), fits on a single
GPU, strong reasoning and tool-calling.

## Step 1: Detect the GPU

```bash
python3 scripts/detect.py
# Remote:
python3 scripts/detect.py --host user@hostname
```

Returns JSON with `gfx_version`, `vram_gb`, `gpu_count`, `rocm_version`.

| gfx_version | Hardware | VRAM |
|---|---|---|
| gfx950 | MI350X / MI355X | 288 GB HBM3E |
| gfx942 | MI300X (192 GB) / MI325X (256 GB) / MI300A (128 GB) | varies |

If `gfx_version` is `unknown`: `amd-smi` ran but found no GPU. Check
`lsmod | grep amdgpu`.

## Step 2: Validate the environment

```bash
python3 scripts/validate.py --auto-fix
# Remote:
python3 scripts/validate.py --auto-fix --host user@hostname
```

Returns JSON with `ready` (bool), `errors`, `warnings`, `fixes_applied`.
Do not proceed if `ready` is `false`.

## Step 3: Refresh recipes (if stale)

Check `fetched_at` in `data/recipes_cache.json`. If older than 24 hours or
the file is missing, refresh:

```bash
python3 scripts/sync_recipes.py
```

This shallow-clones vllm-project/recipes from GitHub and fetches the latest
Docker tag from Docker Hub. Takes ~10 seconds. If it fails, the existing
cache still works.

## Step 4: Construct the Docker command

Read `data/recipes_cache.json` and `data/gpu_overrides.json` directly.
Build the Docker command by combining:

1. **Docker flags** from `gpu_overrides.json > docker_flags` (mandatory for all AMD GPUs)
2. **HF cache mount**: `-v ~/.cache/huggingface:/root/.cache/huggingface`
   (if a shared model cache directory exists on the host, check whether
   `models--*` directories are at the cache root or inside a `hub/`
   subdirectory -- mount accordingly to `/root/.cache/huggingface` or
   `/root/.cache/huggingface/hub`)
3. **Port**: `-p <port>:<port>` (default 8000)
4. **Environment variables**: merge `gpu_configs.<gfx_version>.env_defaults`
   with the recipe's `model.base_env` and `hardware_overrides.amd.extra_env`.
   Always add `--env HF_TOKEN=${HF_TOKEN}`.
5. **Docker image**: use `docker_image` from `recipes_cache.json` top level
   (unless the model needs a pinned image, e.g. GLM-4.5 needs `v0.15.1`).
   If the user specifies a Docker image version, check it against the recipe's
   `model.min_vllm_version`. Warn if the image is older -- the model may crash
   on startup with an opaque "Engine core initialization failed" error.
6. **Model ID**: `--model <HF_ID>`
7. **vLLM args**: combine the recipe's `model.base_args` +
   `hardware_overrides.amd.extra_args` + `features.tool_calling.args` +
   `features.reasoning.args`. Add `--enable-auto-tool-choice` if not present.
   For multi-GPU, add `--tensor-parallel-size N` (see VRAM estimation below).
   For MoE models on multi-GPU, also add `--distributed-executor-backend mp`.
8. **Port arg**: `--port <port>`

If the exact model ID is not in `recipes_cache.json`, check for a base model
match by stripping date/version suffixes (e.g., `Kimi-K2-Instruct` matches
`Kimi-K2-Instruct-0905`). Use the base model's recipe if found.

If no recipe match, check `legacy_models` in `gpu_overrides.json`. If not
there either, use a generic config with
`--enable-auto-tool-choice --trust-remote-code --tool-call-parser hermes`.

**Precision variant selection:** Recipes may offer variants (default, fp8,
nvfp4). Check `gpu_configs.<gfx_version>.precision.native` in
`gpu_overrides.json` before selecting a variant. On gfx942 (MI300X), only
`bf16`, `fp16`, `fp8_fnuz`, and `int8` are hardware-native. MXFP4 and NVFP4
compute is emulated (dequant to BF16 during matmul), but weights stay
compressed in VRAM so quantized models still fit in less memory.
On gfx950 (MI350X), MXFP4 is hardware-native.

**VRAM estimation and fit check:** Before constructing the Docker command,
estimate whether the model fits the available hardware:
```bash
python3 scripts/estimate_vram.py --model-id <HF_ID> --vram-gb <per_gpu_vram> --tp <N>
```
This queries the HuggingFace Hub API (no model download) and returns JSON with:
- `weight_memory_gb` -- total weight size
- `kv_cache_bytes_per_token` -- KV cache cost per token at BF16
- `fit.weights_fit` -- whether weights fit at the given TP
- `fit.recommended_max_model_len` -- max context the GPU can serve
- `fit.context_limited` -- true if KV cache limits context below the
  model's native max
- `fit.min_tp_required` -- minimum TP needed (only if weights don't fit)

**Understanding the overhead:** The script reserves ~4 GB for vLLM's runtime
overhead (activation profiling, HIP graph capture, internal buffers). During
startup, vLLM runs a profiling forward pass to measure peak activations, then
captures HIP graphs for optimized decode. This startup peak is higher than
steady-state. The `remaining_for_kv_gb` field reflects what's left after
weights and this overhead.

Use `remaining_for_kv_gb` to decide:

1. **`remaining_for_kv_gb >= 6`**: safe to run. If `context_limited: true`,
   add `--max-model-len <recommended_max_model_len>` to the vLLM args.
   Mention the FP8 KV cache option (`--kv-cache-dtype fp8`) if the user
   needs longer context (`fit.max_seq_len_fp8_kv` shows the gain).
2. **`remaining_for_kv_gb` between 2 and 6**: tight but worth trying. Launch
   normally. If vLLM OOMs during HIP graph capture (check container logs for
   "out of memory" after "capturing CUDA/HIP graphs"), retry with
   `--enforce-eager` added to the vLLM args. This skips graph capture and
   frees 1-2 GB. The only cost is slightly higher decode latency.
3. **`remaining_for_kv_gb < 2`**: too tight. Will likely OOM during the
   activation profiling step. Do not attempt.
4. **`weights_fit: false` with multiple GPUs**: re-run with
   `--tp <min_tp_required>` and check again.
5. **`weights_fit: false`, not enough GPUs**: look for quantized
   alternatives in this order:
   a. **Recipe variants**: the recipe may have `fp8` or `mxfp4` variants
      with a different `model_id` that points to a quantized checkpoint.
   b. **Same provider**: many providers release quantized versions alongside
      the base model (e.g. `Qwen/Qwen3.5-122B-FP8` from Qwen). Search
      HuggingFace for `<provider>/<model-name>` with FP8/GPTQ/AWQ suffixes.
   c. **AMD quantized**: AMD's Quark team publishes quantized models under
      the `amd/` org on HuggingFace (e.g. `amd/Kimi-K2-Instruct-w-mxfp4-a-fp8`).
      Search for `amd/<model-name>` variants.
   Run `estimate_vram.py` on the quantized model ID to verify it fits,
   then use that model ID instead.
6. **Still doesn't fit**: tell the user the model requires more VRAM than
   available and suggest either a smaller model or multi-GPU hardware.
   Do not attempt to launch.

Docker command template:
```
docker run -d --name vllm-<model-slug> \
  <docker_flags> \
  -v <hf_cache_mount> \
  -p <port>:<port> \
  --env <key>=<value> (for each env var) \
  --env HF_TOKEN=${HF_TOKEN} \
  <docker_image> \
  --model <model_id> \
  <vllm_args> \
  --port <port>
```

## Step 5: Confirm with the user

Before launching, present a summary and ask the user to confirm:
- **Model**: full HuggingFace ID (e.g. `Qwen/Qwen3.5-122B-Instruct`)
- **Precision**: variant being used (e.g. BF16, FP8) and why
- **Weight memory**: from estimate_vram.py
- **GPU**: detected hardware and VRAM
- **TP**: tensor parallelism degree (1, 2, 4, 8)
- **Context**: max achievable context length (and whether it's limited)
- **Port**: which port the endpoint will be on

If a quantized alternative was selected (Step 4 fit check), explain that
the original model doesn't fit and which alternative is being used.

Wait for the user's confirmation before proceeding.

## Step 6: Launch and verify

Before launching, check for port conflicts:
```bash
ss -tlnp 2>/dev/null | grep ':<port> '
```
If a Docker container is on that port, stop it with `docker rm -f <name>`.

Run the Docker command. Then poll health using this loop:

```bash
while docker inspect --format='{{.State.Running}}' <container_name> 2>/dev/null | grep -q true; do
  curl -sf http://localhost:<port>/health && echo "READY" && exit 0
  sleep 60
done
echo "FAILED -- container exited"
```

A 503 during loading is normal. Choose the polling strategy based on
model size (weight memory from hf-mem):

- **Small models (< 100 GB weights)**: run the poll as a blocking command
  with the Bash tool's `timeout` set to 600000 (10 minutes). Most cached
  models are ready within 2-5 minutes.
- **Large models (>= 100 GB weights)**: run the poll with the Bash tool's
  `run_in_background` set to `true`. Then use `TaskOutput` with
  `block: true` and `timeout: 600000` to wait up to 10 minutes per check.
  If the task is still running after that, call `TaskOutput` again with
  the same parameters. This uses only 1 turn per 10-minute wait instead
  of burning a turn every check. The background loop runs until the
  container is healthy or dies.

After health returns 200, send a warmup request (triggers HIP kernel compilation,
~40-45 seconds on gfx942):
```bash
curl -s http://localhost:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model_id>","messages":[{"role":"user","content":"say hi"}],"max_tokens":5}'
```

After the warmup succeeds, present a connection table so the user can call
the endpoint immediately:

| Field | Value |
|-------|-------|
| Model | `<model_id>` |
| Served model name | `<served-model-name or model_id>` |
| Base URL | `http://<host>:<port>/v1` |
| API key | none (local) |
| Port | `<port>` |
| Tensor parallel | `<tp>` |
| Max context | `<context>` |
| GPU | `<detected GPU>` |

Then give a ready-to-run example using those exact values:

```bash
curl -s http://<host>:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model_id>","messages":[{"role":"user","content":"Hello"}]}'
```

## Remote vs. local

All scripts accept `--host user@hostname`. When given, they SSH to the target.
Set `ROCM_SSH_HOST` and `ROCM_SSH_USER` env vars to avoid passing `--host`
every time.

For remote Docker commands, run them over SSH:
```bash
ssh user@host 'docker run -d ...'
```
Use `localhost` for health/warmup curl URLs (curl runs on the remote host).

## Gotchas

**`CUDA_VISIBLE_DEVICES` set to empty string** -- ROCm maps this variable to
`HIP_VISIBLE_DEVICES`. Setting it to an empty string hides all GPUs.
`CUDA_VISIBLE_DEVICES=0,1` works fine for restricting GPUs (same as
`HIP_VISIBLE_DEVICES=0,1`). If the host has it set to empty, unset it:
`unset CUDA_VISIBLE_DEVICES`. Do not pass `--env CUDA_VISIBLE_DEVICES=` (empty)
into Docker -- that also hides all GPUs inside the container.

**FP4BMM crash on gfx942 (MI300X)** -- If the container exits immediately
with a segfault or illegal instruction: `VLLM_ROCM_USE_AITER_FP4BMM` must be
`0` on gfx942. This is set correctly in `gpu_overrides.json` for gfx942.
See vLLM issue #34641.

**`HIP error: no kernel image`** -- The Docker image has no compiled kernel
for your GPU's gfx version. Use `vllm/vllm-openai-rocm:latest`; it includes
gfx942 and gfx950 kernels.

**MLA models need `--block-size 1`** -- DeepSeek-R1/V3, Kimi-K2.5.
Without it the MLA attention backend silently falls back to a slower path.
This is in the recipe args for these models.

**MoE models on multi-GPU need `--distributed-executor-backend mp`** --
Qwen3-235B, GLM-4.5, MiniMax-M2. The default distributed executor does not
work reliably with MoE on ROCm.

**OOM during HIP graph capture** -- If the container logs show "out of memory"
after "capturing CUDA graphs" or "capturing HIP graphs", the model fits in
VRAM but there isn't enough headroom for graph capture. Retry with
`--enforce-eager` added to the vLLM args. This disables graph capture and
frees 1-2 GB. Trade-off: slightly higher decode latency, but the model runs.

**"Engine core initialization failed"** -- This opaque error means the engine
core subprocess died. Check early container logs: `docker logs <name> 2>&1 |
head -50`. Common causes: gated model access denied (license not accepted on
HF), unsupported architecture on this vLLM version, OOM during weight loading,
missing `--trust-remote-code` for custom architectures, or vLLM version too old
for the model (check `min_vllm_version` in the recipe).

**`/dev/kfd` permission denied** -- User is not in the `video` or `render`
group. Fix: `sudo usermod -aG video,render $USER` (requires re-login).

**SSH key not configured** -- The scripts use `BatchMode=yes` SSH. If SSH
fails with `Permission denied (publickey)`, configure key-based access first.

**Restricting GPUs on shared hosts** -- Use `--env HIP_VISIBLE_DEVICES=0,1`
or `--env CUDA_VISIBLE_DEVICES=0,1` to target specific GPUs by index.
`HIP_VISIBLE_DEVICES` is the canonical AMD variable; `CUDA_VISIBLE_DEVICES`
also works (ROCm maps it). Never set either to an empty string.

---

## Reference

Precision compatibility, VRAM estimation, Docker flags, and known quirks:
[reference.md](reference.md)

Referenced files: 12

tracelens-analysis-orchestrator2.84 KB

View saved version →

---
name: tracelens-analysis-orchestrator
description: >-
  Orchestrates modular PyTorch profiler trace analysis with TraceLens: generates perf
  reports, prepares category data, runs system-level and compute-kernel subagents in
  parallel, validates outputs, and writes a prioritized stakeholder report (analysis.md).
  Use when the user asks to follow the analysis orchestrator, run the agentic analysis
  workflow, analyze a trace, compare two traces, or mentions standalone or comparative
  TraceLens analysis.
license: MIT
---

<!--
Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.

See LICENSE for license information.
-->

# Analysis orchestrator

Coordinate **system-level** analysis (CPU/idle, kernel fusion, multi-kernel / comm / memcpy) and **compute-kernel** analysis (GEMM, SDPA, elementwise, etc.): one trace load, shared prep, parallel subagents, then aggregation into `analysis.md`.

## Full procedure

Follow **[reference.md](reference.md)** for every step (user prompts, `<prefix>` / `{CMD}` usage, CLI commands, subagent launch text, validation, report `tee` order, plot embedding, and trace diagnostics).

## Workflow index

```
0. Query User Inputs (Platform, Trace Path(s), Analysis Mode, Environment Setup)
1. Generate Performance Report (branches on analysis mode: training vs inference then, comparison scope)
2-5. Prepare Category Data (GPU Util, Top Ops, Tree Data, Multi-Kernel Data, Category Filtering)
6. System-Level Analysis (PARALLEL) → system_findings/
7. Compute Kernel Subagents (PARALLEL) → category_findings/
   7.5. Aggregate → priority_data.json::findings[]
8. Validate Subagent Outputs
9. load_findings + Model Identification (subagent) → metadata/model_info.json
10. Render performance PNG if agent_extension.py is absent
11. Generate analysis.md (orchestrator writes via <prefix> tee), optional extension, embed PNG
```

## Rules

- **Subagents:** Use the Task tool **only** where reference.md says “subagent” (Steps **6**, **7**, **9**). The orchestrator runs everything else, including Step 7.5, using the command prefix from `<output_dir>/cache/cmd_prefix.txt` (`{CMD}` substitution).
- **Language:** Prefer vendor-agnostic terms (GPU kernels, collective communication, vendor GEMM library, DNN primitives, GPU graph). When quoting trace data, real kernel names are fine.
- **Subagent prompts:** Point each subagent at the checked-in agent file under `TraceLens/Agent/Analysis/skills/analysis-orchestrator/agents/<name>.md` (see reference.md for exact paths and prompt shells).

## Primary outputs

- **Deliverable:** `<output_dir>/analysis.md`
- **Internals:** `system_findings/`, `category_findings/`, `category_data/`, `metadata/`, `perf_report*.xlsx`, CSV folders — see package README for layout.

## Agent layout

Project subagents ship with this skill: `TraceLens/Agent/Analysis/skills/analysis-orchestrator/agents/*.md`.

Referenced files: 21

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 12:00 UTC
Collection status
Collected

plugins_6a57ef89f2d481918513e6133ec3fc18

Download listing JSON