← AMDCONTENT HISTORY

Update to AMD

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.2.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "local-ai-app-integration",
  "description": "Integrates local AI capabilities into applications using Embeddable Lemonade. Use when the user wants to add local AI, offline AI, private AI, on-device AI, a local LLM, local chat, embeddings, image generation, speech-to-text, or text-to-speech to an existing app; replace or supplement OpenAI, Anthropic, Ollama, or other cloud AI APIs with a local backend; only use to convert user apps. Do not use when the user just wants the agent itself to generate images, transcribe, or speak locally in the current workspace, even to cut their own API bill.",
  "included_files": [
    {
      "relative_path": "evals/evals.json",
      "size_in_bytes": 3040
    },
    {
      "relative_path": "evals/files/apikey-guard/main.py",
      "size_in_bytes": 205
    },
    {
      "relative_path": "evals/files/openai-stub/main.py",
      "size_in_bytes": 45
    },
    {
      "relative_path": "reference.md",
      "size_in_bytes": 13905
    },
    {
      "relative_path": "skill-card.md",
      "size_in_bytes": 159
    }
  ],
  "skill_md_contents": "---\nname: local-ai-app-integration\ndescription: >-\n  Integrates local AI capabilities into applications using Embeddable Lemonade.\n  Use when the user wants to add local AI, offline AI, private AI, on-device AI,\n  a local LLM, local chat, embeddings, image generation, speech-to-text, or\n  text-to-speech to an existing app; replace or supplement OpenAI, Anthropic, Ollama, or\n  other cloud AI APIs with a local backend; only use to convert user apps. Do not use when\n  the user just wants the agent itself to generate images, transcribe, or speak locally in\n  the current workspace, even to cut their own API bill.\n---\n\n# Local AI App Integration (Embeddable Lemonade)\n\nAdd a local AI mode to an existing app that already talks to a cloud AI API\n(OpenAI, Anthropic, or Ollama-compatible). The app launches `lemond`, the\nEmbeddable Lemonade binary, as a private subprocess and the existing client\ntalks to it on `http://localhost:PORT/api/v1`. The user gets local, private,\nhardware-optimized inference (CPU, AMD iGPU/dGPU, XDNA2 NPU) with no separate\ninstall.\n\n**What you'll end up with:** one new launcher module (~30 lines), three mandatory changes to the existing HTTP client (`base_url`, `api_key`, and a 120-second HTTP timeout), one vendored binary under `vendor/lemonade/`.\n\n## When this skill is the right tool\n\nUse this skill when **all** of the following are true:\n\n- The app already calls a cloud AI service over HTTP (OpenAI Chat Completions,\n  Anthropic Messages, or Ollama).\n- The user wants that AI to run on the end-user's PC, with the AI engine\n  bundled into the app, not as a separate user install.\n- The target platform is Windows x64 or Linux x64 (macOS embeddable is in beta).\n\nIf the user instead wants a **system-wide** Lemonade Server (one install,\nshared across apps), do not use this skill; point them at\n`https://lemonade-server.ai/install_options.html` and the standard OpenAI base\nURL `http://localhost:13305/api/v1`.\n\n## The opinionated path\n\nThis skill follows one fixed sequence. Do not deviate without a stated reason.\n\n```\n[ ] 1. Survey the app's current AI integration\n[ ] 2. Pick a model + backend profile\n[ ] 3. Place Embeddable Lemonade in the app's tree (full package, not just the binary)\n[ ] 4. Add a `lemond` launcher (subprocess + API key + port + per-stage logging)\n[ ] 5. Re-point the existing client at lemond (base_url, api_key, 120s timeout — all three required)\n[ ] 6. Wait for /api/v1/health, install backend, then PULL the model before first use\n[ ] 7. Wire shutdown and error recovery\n```\n\nTrack progress against this checklist. Move on only when each step verifies.\n\n> **Log every stage.** A local integration has many silent failure points —\n> spawn, health, backend install, model download, first inference. Without a\n> log line at each transition, \"nothing happened\" is indistinguishable from\n> \"broke at stage 3.\" Emit one clear line per stage as you build (see\n> [Step 4](#step-4-add-a-lemond-launcher)); the most common dead-end in this\n> integration — a blank result with no error — is invisible without them.\n\n---\n\n## Step 1: Survey the app\n\nFind every place the app currently calls a cloud AI API. Search the repo for:\n\n- `openai`, `OpenAI(`, `chat.completions`, `responses.create`\n- `anthropic`, `Anthropic(`, `messages.create`\n- `api.openai.com`, `api.anthropic.com`, `localhost:11434` (Ollama)\n- `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`\n\nRecord three things before continuing:\n\n1. **Client library and language** (e.g., `openai-python`, `openai-node`,\n   `@anthropic-ai/sdk`, `go-openai`, raw `fetch`).\n2. **Modalities used:** text chat, tool calling, embeddings, image gen,\n   transcription, TTS. This drives the model + backend choice in Step 2.\n3. **One single place** where the base URL and API key are constructed. If\n   there isn't one, refactor to one before going further. Local-mode toggling\n   must flip exactly one config object.\n4. **Any API-key gating** that blocks the app before a key is entered\n   (onboarding walls, validators that reject empty keys, startup checks that\n   disable AI until a key exists). Note each one — Step 5 bypasses them in\n   local mode.\n\n## Step 2: Pick a model + backend profile\n\nChoose **one** default profile based on the app's primary modality. Do not\nship a buffet. Ship one good default and document how the user can override\nit.\n\n| App's primary need | Default model | Recipe | Why |\n|---|---|---|---|\n| General chat / assistant | `Qwen3-4B-GGUF` | `llamacpp` | Small, fast, good tool calling, fits 8GB systems |\n| Coding assistant | `Qwen2.5-Coder-7B-Instruct-GGUF` | `llamacpp` | Strong code, runs on iGPU |\n| Vision / multimodal chat | `Gemma-4-E2B-it-GGUF` | `llamacpp` | Small multimodal default |\n| NPU-first on Ryzen AI | `Llama-3.2-3B-Instruct-Hybrid` | `ryzenai-llm` | XDNA2 NPU on Windows |\n| Speech-to-text (Windows) | `Whisper-Large-v3-Turbo` | `whispercpp` | One model; probe picks NPU → iGPU/dGPU → CPU automatically |\n| Speech-to-text (Linux NPU) | `whisper-v3-turbo-FLM` | `flm` | Linux NPU path; falls back to `whispercpp` iGPU/CPU off-NPU |\n| Text-to-speech | `kokoro-v1` | `kokoro` | CPU-only, low latency |\n| Image generation | `SDXL-Turbo` | `sd-cpp` | Single-step generation |\n\nFor the LLM backend, default to `llamacpp` and let `lemond` pick\n`rocm` → `vulkan` → `cpu` automatically by leaving `llamacpp_backend`\nunset. Override only if the app has hard hardware requirements.\n\n**Scope: this skill selects a backend once at integration time on the\ndeveloper's machine.** Runtime fallback based on the end user's hardware is\nout of scope. Bundle `vulkan` as the universal fallback so the app works on\nany machine. If the dev machine has an NPU and the chosen recipe supports it,\nthe skill will use the NPU backend — otherwise it falls back to `vulkan`.\n\n> **Note:** having an NPU does not mean every recipe supports NPU. Confirm\n> the recipe/backend pair is `installed` or `installable` via\n> `GET /api/v1/system-info` before committing to it. See\n> [reference.md](reference.md#hardware-probing-with-v1system-info) for\n> per-recipe decision rules.\n\nFor more options and tradeoffs, see [reference.md](reference.md).\n\n## Step 3: Place Embeddable Lemonade in the app's tree and install backends\n\n**Get the embeddable artifact** from the latest Lemonade release:\n\n```\nhttps://github.com/lemonade-sdk/lemonade/releases/latest\n```\n\nDownload the file matching your target OS:\n\n- Windows: `lemonade-embeddable-{VERSION}-windows-x64.zip`\n- Linux:   `lemonade-embeddable-{VERSION}-ubuntu-x64.tar.gz`\n\n> **Don't hand-build the download URL from the tag.** The git tag carries a\n> leading `v` (e.g. `v10.8.0`) but the asset filename strips it\n> (`lemonade-embeddable-10.8.0-...`), so using the tag verbatim 404s. Ask the\n> GitHub API for the asset by its stable name pattern and use the URL it\n> returns, as below — this stays correct across version and naming changes.\n\n**First, create the target directory** — it does not exist in a fresh repo:\n\n```powershell\n# Windows\nNew-Item -ItemType Directory -Force vendor\\lemonade\n```\n\n```bash\n# Linux\nmkdir -p vendor/lemonade\n```\n\nThen download and unpack on Windows (PowerShell):\n\n```powershell\n$rel = Invoke-RestMethod https://api.github.com/repos/lemonade-sdk/lemonade/releases/latest\n$asset = $rel.assets | Where-Object { $_.name -like \"lemonade-embeddable-*-windows-x64.zip\" } | Select-Object -First 1\nInvoke-WebRequest $asset.browser_download_url -OutFile lemond.zip\nExpand-Archive lemond.zip -DestinationPath \"$env:TEMP\\lemond-unpack\"\n$folder = $asset.name -replace '\\.zip$',''   # unpacked dir = asset name without .zip\nCopy-Item -Recurse \"$env:TEMP\\lemond-unpack\\$folder\\*\" vendor\\lemonade\\\n# Sanity check: resources/ must be nested under vendor\\lemonade\\ (not flattened)\nif (-not (Test-Path vendor\\lemonade\\resources\\*.json)) { throw \"resources/ missing — re-extract and copy again\" }\n```\n\nOn Linux (bash):\n\n```bash\nURL=$(curl -s https://api.github.com/repos/lemonade-sdk/lemonade/releases/latest \\\n  | grep browser_download_url | grep ubuntu-x64.tar.gz | cut -d'\"' -f4)\ncurl -L \"$URL\" | tar -xz --strip-components=1 -C vendor/lemonade\n```\n\n> **Copy the full package, not just the binary.** The archive contains\n> `lemond[.exe]`, `lemonade[.exe]`, `LICENSE`, and `resources/`. The\n> `resources/` directory is required — without it lemond starts and passes the\n> health check but fails on every model and backend request. Copying only the\n> binary produces a server that looks healthy but cannot function.\n\n> **`lemond` vs `lemonade` CLI:** `lemond` is the embedded server binary that\n> ships with the app. The `lemonade` CLI is a separate packaging tool used\n> only during development/build time to install backends. The same embeddable\n> archive unpacked above already contains a matching `lemonade[.exe]` next to\n> `lemond[.exe]`, so its version aligns with the bundled `lemond`. Do **not**\n> `pip install lemonade-sdk` to get it: the PyPI package is a separate, older\n> release line whose ports, model names, and install API do not match the\n> `lemond` bundled here, and mixing the two is a known source of silent\n> version mismatches. Keep the `lemonade` CLI, `lemond`, and the backends all\n> from the one release downloaded in this step so their versions stay aligned.\n\nThe expected layout **after setup** (first run + backend install). A freshly\nunzipped package contains only `lemond[.exe]`, `lemonade[.exe]`, `LICENSE`, and\n`resources/` — the items below are created later, as their comments note:\n\n```\nvendor/lemonade/\n  lemond[.exe]                     # the only binary the app ships\n  LICENSE\n  config.json                      # generated on first run; commit a seed copy\n  resources/\n    server_models.json             # do not edit; use GET /api/v1/models at runtime\n    backend_versions.json\n  bin/                             # backends bundled at packaging time\n    llamacpp/vulkan/llama-server[.exe]\n  models/                          # pre-bundled model weights (optional)\n    models--unsloth--Qwen3-4B-GGUF/\n```\n\n> **`server_models.json`:** Do not edit or rely on this file. It can be stale.\n> The only authoritative model list is `GET /api/v1/models` on a running\n> `lemond` instance with the backend already installed.\n\n**Bundle decisions: pick deliberately**\n\n- **Backends:** Bundle `llamacpp:vulkan` at packaging time (works on every\n  GPU). Install `llamacpp:rocm` at first run on supported AMD systems via\n  `POST /api/v1/install` after probing `GET /api/v1/system-info`. Never ship\n  every backend, or the artifact balloons.\n- **Models:** Either bundle the default model under `models/` (offline\n  install, larger installer) **or** pull on first run with\n  `POST /api/v1/pull` (smaller installer, needs network). Pick one and\n  document it.\n- **`models_dir`:** Set to `./models` in `config.json` to keep weights\n  private to the app. Leave as `auto` only if the user explicitly wants to\n  share weights with other apps.\n\n**Backend install timing — two distinct paths:**\n\n> **Packaging time** (developer machine, before bundling). Use the lemonade\n> CLI that shipped inside `vendor/lemonade/` so it matches the bundled\n> `lemond` version (prefix with `./` or the full path):\n> ```\n> vendor/lemonade/lemonade backends install llamacpp:vulkan\n> vendor/lemonade/lemonade backends install flm:npu    # Windows NPU path only\n> ```\n> This bakes the backend binaries into `vendor/lemonade/bin/` before the app\n> ships. `lemond` does not need to be running. Use a modern `lemonade` CLI\n> whose version matches the bundled `lemond` (the copy in the archive you\n> unpacked works); do not `pip install lemonade-sdk` for it.\n>\n> **First-run / runtime** (user's machine, after `lemond` is running):\n> ```http\n> POST /api/v1/install\n> {\"recipe\": \"llamacpp\", \"backend\": \"rocm\"}\n> ```\n> Use this for hardware-specific backends (e.g. `llamacpp:rocm`) that cannot\n> be bundled universally. `lemond` must already be running (Step 4 complete).\n\n## Step 4: Add a `lemond` launcher\n\nWrite the launcher as a new module named **`lemond_launcher.py`** (or\n`lemond_launcher.<ext>` for the app's language). It is a thin process\nsupervisor. Its only jobs:\n\n1. Generate a fresh random API key: `key = secrets.token_urlsafe(32)`\n2. Pick a free localhost port: bind a `socket` to port 0, read back the assigned port, close it.\n3. Spawn lemond as a `subprocess`: `subprocess.Popen([LEMOND_BIN, LEMOND_DIR, \"--port\", str(port)], env={**os.environ, \"LEMONADE_API_KEY\": key})`\n4. Poll `GET /api/v1/health` with `Authorization: Bearer {key}` in a loop until HTTP 200 — this is the only correct readiness check.\n5. Expose the chosen `port` and `key` to the rest of the app.\n\n> **Log one line per lifecycle stage.** Build the logging in from the start —\n> not as an afterthought when something breaks. Each silent transition needs a\n> visible marker so a failure points at the exact stage. Aim for:\n>\n> ```\n> [lemond] Starting on port <port>\n> [lemond] Healthy on port <port>\n> [lemond] <recipe>:<backend> installed        (or: already installed / install failed)\n> [lemond] Pulling model <name>...             then: Model <name> ready  (or: pull returned <status>)\n> [local]  <modality> result: <value>          (first inference output — empty string here = unpulled model)\n> ```\n>\n> Logging the **first inference result verbatim** is what turns the\n> silent-empty failure (Step 6) from a multi-hour mystery into a one-line\n> diagnosis. Route these through the app's normal logging so they can be quieted\n> for release.\n\n> **Dev-mode file watchers:** If the app runs with a file watcher (Tauri,\n> Electron, Next.js, Vite, etc.) that watches the source tree, ensure\n> `vendor/lemonade/` is excluded from the watched paths. Lemond writes config\n> and cache files at runtime; a watcher that picks these up will restart the\n> app, kill the lemond subprocess, and spawn a new one on a new port —\n> silently breaking any in-flight transcription. Add `vendor/` (or the\n> equivalent) to the watcher's ignore list before testing.\n\n**Use the reference implementation from [reference.md § Reference launchers](reference.md#reference-launchers) directly** — copy it verbatim and adapt only the `LEMOND_DIR` path. Do not write a launcher from scratch. The reference Python launcher uses `secrets` (for the API key), `socket` (for the free-port probe), and `subprocess` (to spawn lemond); the Node.js launcher uses the equivalent stdlib modules. Both handle port-race retries and health polling correctly.\n\nReadiness is always determined by polling the exact endpoint\n`GET http://127.0.0.1:<port>/api/v1/health` and checking for HTTP 200 — never\nby reading `lemond`'s stdout or stderr. Any health-check helper you write must\nhit that `/api/v1/health` path.\n\n## Step 5: Re-point the existing client at `lemond`\n\nMake **three** changes to the app's existing client construction — all three\nare required, not optional:\n\n1. Set `base_url` to `http://127.0.0.1:{port}/api/v1`\n2. Set `api_key` to the launcher key\n3. **Set the HTTP timeout to 120 seconds** — this is mandatory, not optional\n\nThe 120-second timeout is not a tuning suggestion. The default on most HTTP\nclients is 30s, which is shorter than lemond's first-run model load time on\nreal hardware. Without it the request silently times out and the UI shows\nnothing, which is indistinguishable from a broken integration.\n\n**Python (openai) — the exact change to make:**\n\n```python\nimport httpx\nfrom openai import OpenAI\n\nproc, key, port = start_lemond()\nclient = OpenAI(\n    base_url=f\"http://127.0.0.1:{port}/api/v1\",\n    api_key=key,\n    http_client=httpx.Client(timeout=120),  # required: 120s for first-run model load\n)\n```\n\nFor other clients:\n\n| Existing client | New `base_url` | New auth | Timeout |\n|---|---|---|---|\n| `openai-python` | `http://127.0.0.1:{port}/api/v1` | `api_key=key` | `httpx.Client(timeout=120)` |\n| `openai-node` | `http://127.0.0.1:{port}/api/v1` | `apiKey: key` | `timeout: 120000` |\n| `@anthropic-ai/sdk` | `http://127.0.0.1:{port}/api/v1` | `apiKey: key` | `timeout: 120000` |\n| Raw `fetch` / `requests` | same | `Authorization: Bearer {key}` | set per-request |\n| Ollama-compatible code | `http://127.0.0.1:{port}/api/v0` | pass key anyway | 120s |\n\nThe model identifier on requests stays a Lemonade model name (e.g.\n`Qwen3-4B-GGUF`), not the cloud name.\n\n**Local mode needs no cloud API key — at all.** This is a defining property of\nlocal mode, not an edge case: there is no cloud service to authenticate to, so\nnothing should ever ask the user for a key. Any onboarding wall, validator, or\nstartup check that demands one must not block local-mode users. Concretely:\n\n- Skip or auto-satisfy the key-entry screen in local mode.\n- Treat local mode as already-authorized in every validation path — an\n  empty-key check must short-circuit to \"valid\" when the active mode is local,\n  never throw \"API key not configured\".\n- Re-enable the gate **only** for cloud mode.\n\nThe `lemond` key from Step 4 is generated internally by the launcher and used\nonly for the local loopback connection, so the user never sees or enters one;\nany UI placeholder (e.g. `\"local\"`) is fine. Flipping into local mode should\nnever strand the user on a key-entry wall.\n\n## Step 6: Health, backend, then pull the model — *before* first inference\n\n`GET /api/v1/health` returning 200 means the **server** is up. It does **not**\nmean inference will work. Before the first real request succeeds, three more\nthings must be true: the backend for your modality is installed, the model's\nweights are **downloaded to disk**, and (on the first call) the model is loaded\ninto memory. Treating health=200 as \"ready\" is the single biggest cause of a\nbroken-looking integration.\n\n**Do not call `POST /api/v1/load` at startup.** Lemond lazy-loads the model\ninto memory on the first inference request and handles that step on its own.\nPre-loading is unreliable across lemond versions (the `/load` request body\nshape has changed between releases) and a malformed call can crash or\ndestabilise the server before the user takes any action. Loading is the one\nstep you let lemond do lazily — pulling is not.\n\n### Pull the model so it exists on disk\n\nLazy-load only loads weights that are **already downloaded**. If the model was\nnever pulled, the first inference does not error — lemond returns an empty /\nblank result with HTTP 200. So after health passes and the backend is\ninstalled, proactively pull the model:\n\n```http\nPOST /api/v1/pull\n{\"model\": \"Whisper-Large-v3-Turbo\"}\n```\n\nThis is **idempotent** — a no-op if the weights are already present, a download\nif they are not. Run it once during setup (after backend install, before the\nfirst user-triggered inference) and log the result.\n\n- **Default model** (the one you chose in Step 2): pull it by name as above.\n- **Custom / user-overridden model:** do not assume it exists. Confirm it is a\n  real Lemonade model first via `GET /api/v1/models` (the **only** trusted\n  catalog — see [reference.md](reference.md)), then pull it the same way. A\n  model appearing in the catalog is **not** proof its weights are downloaded;\n  a successful pull is.\n\n> **Silent-empty is almost always an unpulled model.** If inference returns an\n> empty string / blank output with no HTTP error, the model was not downloaded.\n> Check your pull step before debugging anything else — this is the failure mode\n> that wastes the most time. Log the pull result and the first inference result\n> (see Step 4) so this is diagnosable from the console, not by guesswork.\n\n### Surface the *whole* setup, not just model load\n\nFirst-run cold start is more than a model load. The full sequence is:\n\n```\nserver spawn  →  health 200  →  backend install  →  model download  →  model load  →  first result\n```\n\nOn a fresh machine, backend install and model download can each take from tens\nof seconds to several **minutes** (multi-GB weights over the network). Model\nload alone is 10–30s. An app that shows nothing during this will look frozen.\n\nMinimum: show a loading indicator or status message (\"Setting up local AI…\")\nfrom the moment setup begins until the first response arrives — covering the\n*entire* sequence above, not just the final load. The simplest implementation\nis a flag set when setup/first-request starts and cleared when the first\nresponse arrives. Once the model is pulled and loaded once, subsequent runs are\nfast; the long wait is first-run only.\n\n## Step 7: Lifecycle and recovery\n\nThese are the only failure modes worth handling. Do not over-engineer.\n\n| Symptom | Cause | Recovery |\n|---|---|---|\n| **Inference returns empty / blank with HTTP 200, no error** | Model never pulled: backend is installed but weights are absent, so lazy-load has nothing to load | `POST /api/v1/pull` with `{\"model\":\"...\"}`, wait for success, retry. Log the pulled result and the first inference result. This is the most common silent failure — see [Step 6](#step-6-health-backend-then-pull-the-model--before-first-inference) |\n| `POST /api/v1/load` returns 404 / model not found | Model not pulled yet (same root cause as the empty-result row above) | `POST /api/v1/pull` with `{\"model\": \"...\"}` then retry `/api/v1/load` |\n| `POST /api/v1/load` returns 500 with backend error | Backend not installed for this hardware | `GET /api/v1/system-info`, pick a supported backend, `POST /api/v1/install` with `{\"recipe\": \"...\", \"backend\": \"...\"}`, retry |\n| Subprocess exits immediately | Port race: another process grabbed the port between `freePort()` and lemond binding | The reference launcher retries with a fresh port automatically (3 attempts) |\n| `/api/v1/health` never returns 200 | First-run backend extraction is slow on cold disk | Extend timeout to 90s on first launch, 30s after |\n| HTTP 401 on every request | Forgot the `Authorization: Bearer` header | Audit the client config because Lemonade rejects unauth'd calls when `LEMONADE_API_KEY` is set |\n\n**Shutdown:** On app exit, `proc.terminate()` (Unix) or\n`proc.kill()` (Windows). `lemond` flushes config and exits cleanly within a\ncouple of seconds. Always wait on the process; never orphan it.\n\n**Do not** parse `lemond` stdout to detect readiness; use the HTTP\n`/api/v1/health` probe. Stdout format is not a stable contract.\n\n---\n\n## Verification checklist\n\nThe integration is done when **all** of these are true:\n\n- [ ] `vendor/lemonade/` contains the full package: `lemond[.exe]`,\n      `lemonade[.exe]`, `LICENSE`, and `resources/` — not just the binary.\n- [ ] `lemond` starts as a subprocess with a fresh API key per launch.\n- [ ] `GET /api/v1/health` returns 200 within the timeout.\n- [ ] The default model is pulled (or bundled) before the first inference; a\n      custom/overridden model is confirmed via `GET /api/v1/models` and then\n      pulled. A blank result with no error means this step was skipped.\n- [ ] Each lifecycle stage logs a clear line (spawn, health, backend install,\n      model pull, first result) so a failure is diagnosable from the console.\n- [ ] The existing client's chat / image / speech call returns a valid\n      response with the base URL and key swapped, with no other code changed.\n- [ ] First-run latency is surfaced: the interface shows a loading state from the\n      moment the first inference request is sent until the response arrives.\n- [ ] The HTTP client timeout is set to 120 seconds.\n- [ ] In local mode the app requires **no** cloud API key: no onboarding wall,\n      validator, or startup check blocks the user, and no code path throws\n      \"API key not configured\" when the active mode is local.\n- [ ] If the app uses a dev-mode file watcher, `vendor/lemonade/` is excluded\n      from the watched paths so runtime writes by lemond do not trigger restarts.\n- [ ] Killing the parent process leaves no `lemond` subprocess behind.\n- [ ] On a fresh machine without the optimal backend, the app still works\n      via the Vulkan fallback bundled in `bin/`.\n\nIf any box is unchecked, do not declare the task complete.\n\n---\n\n## Reference\n\nFor detailed model catalog, backend selection matrix, full endpoint reference,\nconfig keys, and per-model `recipe_options.json` tuning, see\n[reference.md](reference.md).\n"
}

SHA-256: 871c9df9f412055206fe1d7d61440b150e8f24b0f5e73b47273e904467ca2f99