← AMDCONTENT HISTORY

Update to AMD

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.2.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "serving-llms-on-instinct",
  "description": "Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill whenever the user wants to run, serve, deploy, start, host, or launch a language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X. Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM, or asks how to get a model running on AMD data center hardware. Use when the user asks \"run Qwen3\", \"serve DeepSeek\", \"start a vLLM endpoint\", \"get a model running on my AMD machine\", or any similar phrasing. Handles the full flow: GPU detection, environment validation, vLLM configuration, launch, and health verification. Do not use for NVIDIA GPUs, consumer AMD GPUs (RX series, Radeon), Ryzen AI, NPU, MI250X, or MI100.",
  "included_files": [
    {
      "relative_path": "data/blacklist.json",
      "size_in_bytes": 2050
    },
    {
      "relative_path": "data/gpu_overrides.json",
      "size_in_bytes": 5257
    },
    {
      "relative_path": "data/recipes_cache.json",
      "size_in_bytes": 696133
    },
    {
      "relative_path": "evals/evals.json",
      "size_in_bytes": 2488
    },
    {
      "relative_path": "evals/hooks.py",
      "size_in_bytes": 3239
    },
    {
      "relative_path": "evals/machine.yml",
      "size_in_bytes": 22
    },
    {
      "relative_path": "reference.md",
      "size_in_bytes": 4621
    },
    {
      "relative_path": "scripts/detect.py",
      "size_in_bytes": 4636
    },
    {
      "relative_path": "scripts/estimate_vram.py",
      "size_in_bytes": 8845
    },
    {
      "relative_path": "scripts/sync_recipes.py",
      "size_in_bytes": 6337
    },
    {
      "relative_path": "scripts/validate.py",
      "size_in_bytes": 8453
    },
    {
      "relative_path": "skill-card.md",
      "size_in_bytes": 244
    }
  ],
  "skill_md_contents": "---\nname: serving-llms-on-instinct\ndescription: >-\n  Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill\n  whenever the user wants to run, serve, deploy, start, host, or launch a\n  language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X.\n  Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM,\n  or asks how to get a model running on AMD data center hardware. Use when the\n  user asks \"run Qwen3\", \"serve DeepSeek\", \"start a vLLM endpoint\", \"get a\n  model running on my AMD machine\", or any similar phrasing. Handles the full\n  flow: GPU detection, environment validation, vLLM configuration, launch, and\n  health verification. Do not use for NVIDIA GPUs, consumer AMD GPUs (RX\n  series, Radeon), Ryzen AI, NPU, MI250X, or MI100.\nallowed-tools: Bash, Read\n---\n\n# Serving LLMs on AMD Instinct\n\nGet a vLLM endpoint running on AMD Instinct GPU hardware.\n\n## Prerequisites\n\n- ROCm driver and `amd-smi` installed on the GPU host\n- Docker running and accessible (check with `docker ps`)\n- `/dev/kfd` and `/dev/dri` present on the GPU host\n- HuggingFace token in `HF_TOKEN` env var (required for gated models; not\n  required for Qwen3 or Gemma). For gated models (Llama 3.2, Gemma, etc.),\n  the HF token must belong to an account that has accepted the model's license\n  at `huggingface.co/<model_id>`. A valid token without license acceptance will\n  fail with an opaque \"Engine core initialization failed\" error.\n- For remote GPU: SSH key access configured (`ssh <user>@<host>` must work\n  without a password prompt). If only password access is available, set up\n  keys first: `ssh-copy-id <user>@<host>`\n\n## Data files\n\nRead these files directly to get model and GPU configuration:\n\n- **`data/recipes_cache.json`** -- model configs synced from\n  [vllm-project/recipes](https://github.com/vllm-project/recipes). Each entry\n  under `models.<HF_ID>.recipe` contains the full recipe with `model.base_args`,\n  `model.base_env`, `features.tool_calling.args`, `features.reasoning.args`,\n  `hardware_overrides.amd.extra_args`, `hardware_overrides.amd.extra_env`.\n  The top-level `docker_image` field has the latest resolved vLLM ROCm image.\n\n- **`data/gpu_overrides.json`** -- GPU-specific configuration. Contains\n  `docker_flags` (mandatory for all AMD Instinct), `gpu_configs` keyed by\n  gfx_version with `env_defaults` and `workarounds`, and `legacy_models` for\n  models not yet in vLLM recipes.\n\n- **`data/blacklist.json`** -- models in vLLM recipes that cannot be served\n  as LLM endpoints. Includes diffusion/image/audio generation models, embedding\n  models, rerankers, ASR models needing audio pipelines, and models requiring\n  unreleased vLLM nightly builds. Check this before attempting to serve a model.\n  If the user requests a blacklisted model, explain why it won't work and\n  suggest an alternative.\n\nIf the user doesn't specify a model, default to **Qwen/Qwen3.5-9B**: dense\nmultimodal with MTP, Apache 2.0 license (no HF token needed), fits on a single\nGPU, strong reasoning and tool-calling.\n\n## Step 1: Detect the GPU\n\n```bash\npython3 scripts/detect.py\n# Remote:\npython3 scripts/detect.py --host user@hostname\n```\n\nReturns JSON with `gfx_version`, `vram_gb`, `gpu_count`, `rocm_version`.\n\n| gfx_version | Hardware | VRAM |\n|---|---|---|\n| gfx950 | MI350X / MI355X | 288 GB HBM3E |\n| gfx942 | MI300X (192 GB) / MI325X (256 GB) / MI300A (128 GB) | varies |\n\nIf `gfx_version` is `unknown`: `amd-smi` ran but found no GPU. Check\n`lsmod | grep amdgpu`.\n\n## Step 2: Validate the environment\n\n```bash\npython3 scripts/validate.py --auto-fix\n# Remote:\npython3 scripts/validate.py --auto-fix --host user@hostname\n```\n\nReturns JSON with `ready` (bool), `errors`, `warnings`, `fixes_applied`.\nDo not proceed if `ready` is `false`.\n\n## Step 3: Refresh recipes (if stale)\n\nCheck `fetched_at` in `data/recipes_cache.json`. If older than 24 hours or\nthe file is missing, refresh:\n\n```bash\npython3 scripts/sync_recipes.py\n```\n\nThis shallow-clones vllm-project/recipes from GitHub and fetches the latest\nDocker tag from Docker Hub. Takes ~10 seconds. If it fails, the existing\ncache still works.\n\n## Step 4: Construct the Docker command\n\nRead `data/recipes_cache.json` and `data/gpu_overrides.json` directly.\nBuild the Docker command by combining:\n\n1. **Docker flags** from `gpu_overrides.json > docker_flags` (mandatory for all AMD GPUs)\n2. **HF cache mount**: `-v ~/.cache/huggingface:/root/.cache/huggingface`\n   (if a shared model cache directory exists on the host, check whether\n   `models--*` directories are at the cache root or inside a `hub/`\n   subdirectory -- mount accordingly to `/root/.cache/huggingface` or\n   `/root/.cache/huggingface/hub`)\n3. **Port**: `-p <port>:<port>` (default 8000)\n4. **Environment variables**: merge `gpu_configs.<gfx_version>.env_defaults`\n   with the recipe's `model.base_env` and `hardware_overrides.amd.extra_env`.\n   Always add `--env HF_TOKEN=${HF_TOKEN}`.\n5. **Docker image**: use `docker_image` from `recipes_cache.json` top level\n   (unless the model needs a pinned image, e.g. GLM-4.5 needs `v0.15.1`).\n   If the user specifies a Docker image version, check it against the recipe's\n   `model.min_vllm_version`. Warn if the image is older -- the model may crash\n   on startup with an opaque \"Engine core initialization failed\" error.\n6. **Model ID**: `--model <HF_ID>`\n7. **vLLM args**: combine the recipe's `model.base_args` +\n   `hardware_overrides.amd.extra_args` + `features.tool_calling.args` +\n   `features.reasoning.args`. Add `--enable-auto-tool-choice` if not present.\n   For multi-GPU, add `--tensor-parallel-size N` (see VRAM estimation below).\n   For MoE models on multi-GPU, also add `--distributed-executor-backend mp`.\n8. **Port arg**: `--port <port>`\n\nIf the exact model ID is not in `recipes_cache.json`, check for a base model\nmatch by stripping date/version suffixes (e.g., `Kimi-K2-Instruct` matches\n`Kimi-K2-Instruct-0905`). Use the base model's recipe if found.\n\nIf no recipe match, check `legacy_models` in `gpu_overrides.json`. If not\nthere either, use a generic config with\n`--enable-auto-tool-choice --trust-remote-code --tool-call-parser hermes`.\n\n**Precision variant selection:** Recipes may offer variants (default, fp8,\nnvfp4). Check `gpu_configs.<gfx_version>.precision.native` in\n`gpu_overrides.json` before selecting a variant. On gfx942 (MI300X), only\n`bf16`, `fp16`, `fp8_fnuz`, and `int8` are hardware-native. MXFP4 and NVFP4\ncompute is emulated (dequant to BF16 during matmul), but weights stay\ncompressed in VRAM so quantized models still fit in less memory.\nOn gfx950 (MI350X), MXFP4 is hardware-native.\n\n**VRAM estimation and fit check:** Before constructing the Docker command,\nestimate whether the model fits the available hardware:\n```bash\npython3 scripts/estimate_vram.py --model-id <HF_ID> --vram-gb <per_gpu_vram> --tp <N>\n```\nThis queries the HuggingFace Hub API (no model download) and returns JSON with:\n- `weight_memory_gb` -- total weight size\n- `kv_cache_bytes_per_token` -- KV cache cost per token at BF16\n- `fit.weights_fit` -- whether weights fit at the given TP\n- `fit.recommended_max_model_len` -- max context the GPU can serve\n- `fit.context_limited` -- true if KV cache limits context below the\n  model's native max\n- `fit.min_tp_required` -- minimum TP needed (only if weights don't fit)\n\n**Understanding the overhead:** The script reserves ~4 GB for vLLM's runtime\noverhead (activation profiling, HIP graph capture, internal buffers). During\nstartup, vLLM runs a profiling forward pass to measure peak activations, then\ncaptures HIP graphs for optimized decode. This startup peak is higher than\nsteady-state. The `remaining_for_kv_gb` field reflects what's left after\nweights and this overhead.\n\nUse `remaining_for_kv_gb` to decide:\n\n1. **`remaining_for_kv_gb >= 6`**: safe to run. If `context_limited: true`,\n   add `--max-model-len <recommended_max_model_len>` to the vLLM args.\n   Mention the FP8 KV cache option (`--kv-cache-dtype fp8`) if the user\n   needs longer context (`fit.max_seq_len_fp8_kv` shows the gain).\n2. **`remaining_for_kv_gb` between 2 and 6**: tight but worth trying. Launch\n   normally. If vLLM OOMs during HIP graph capture (check container logs for\n   \"out of memory\" after \"capturing CUDA/HIP graphs\"), retry with\n   `--enforce-eager` added to the vLLM args. This skips graph capture and\n   frees 1-2 GB. The only cost is slightly higher decode latency.\n3. **`remaining_for_kv_gb < 2`**: too tight. Will likely OOM during the\n   activation profiling step. Do not attempt.\n4. **`weights_fit: false` with multiple GPUs**: re-run with\n   `--tp <min_tp_required>` and check again.\n5. **`weights_fit: false`, not enough GPUs**: look for quantized\n   alternatives in this order:\n   a. **Recipe variants**: the recipe may have `fp8` or `mxfp4` variants\n      with a different `model_id` that points to a quantized checkpoint.\n   b. **Same provider**: many providers release quantized versions alongside\n      the base model (e.g. `Qwen/Qwen3.5-122B-FP8` from Qwen). Search\n      HuggingFace for `<provider>/<model-name>` with FP8/GPTQ/AWQ suffixes.\n   c. **AMD quantized**: AMD's Quark team publishes quantized models under\n      the `amd/` org on HuggingFace (e.g. `amd/Kimi-K2-Instruct-w-mxfp4-a-fp8`).\n      Search for `amd/<model-name>` variants.\n   Run `estimate_vram.py` on the quantized model ID to verify it fits,\n   then use that model ID instead.\n6. **Still doesn't fit**: tell the user the model requires more VRAM than\n   available and suggest either a smaller model or multi-GPU hardware.\n   Do not attempt to launch.\n\nDocker command template:\n```\ndocker run -d --name vllm-<model-slug> \\\n  <docker_flags> \\\n  -v <hf_cache_mount> \\\n  -p <port>:<port> \\\n  --env <key>=<value> (for each env var) \\\n  --env HF_TOKEN=${HF_TOKEN} \\\n  <docker_image> \\\n  --model <model_id> \\\n  <vllm_args> \\\n  --port <port>\n```\n\n## Step 5: Confirm with the user\n\nBefore launching, present a summary and ask the user to confirm:\n- **Model**: full HuggingFace ID (e.g. `Qwen/Qwen3.5-122B-Instruct`)\n- **Precision**: variant being used (e.g. BF16, FP8) and why\n- **Weight memory**: from estimate_vram.py\n- **GPU**: detected hardware and VRAM\n- **TP**: tensor parallelism degree (1, 2, 4, 8)\n- **Context**: max achievable context length (and whether it's limited)\n- **Port**: which port the endpoint will be on\n\nIf a quantized alternative was selected (Step 4 fit check), explain that\nthe original model doesn't fit and which alternative is being used.\n\nWait for the user's confirmation before proceeding.\n\n## Step 6: Launch and verify\n\nBefore launching, check for port conflicts:\n```bash\nss -tlnp 2>/dev/null | grep ':<port> '\n```\nIf a Docker container is on that port, stop it with `docker rm -f <name>`.\n\nRun the Docker command. Then poll health using this loop:\n\n```bash\nwhile docker inspect --format='{{.State.Running}}' <container_name> 2>/dev/null | grep -q true; do\n  curl -sf http://localhost:<port>/health && echo \"READY\" && exit 0\n  sleep 60\ndone\necho \"FAILED -- container exited\"\n```\n\nA 503 during loading is normal. Choose the polling strategy based on\nmodel size (weight memory from hf-mem):\n\n- **Small models (< 100 GB weights)**: run the poll as a blocking command\n  with the Bash tool's `timeout` set to 600000 (10 minutes). Most cached\n  models are ready within 2-5 minutes.\n- **Large models (>= 100 GB weights)**: run the poll with the Bash tool's\n  `run_in_background` set to `true`. Then use `TaskOutput` with\n  `block: true` and `timeout: 600000` to wait up to 10 minutes per check.\n  If the task is still running after that, call `TaskOutput` again with\n  the same parameters. This uses only 1 turn per 10-minute wait instead\n  of burning a turn every check. The background loop runs until the\n  container is healthy or dies.\n\nAfter health returns 200, send a warmup request (triggers HIP kernel compilation,\n~40-45 seconds on gfx942):\n```bash\ncurl -s http://localhost:<port>/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"<model_id>\",\"messages\":[{\"role\":\"user\",\"content\":\"say hi\"}],\"max_tokens\":5}'\n```\n\nAfter the warmup succeeds, present a connection table so the user can call\nthe endpoint immediately:\n\n| Field | Value |\n|-------|-------|\n| Model | `<model_id>` |\n| Served model name | `<served-model-name or model_id>` |\n| Base URL | `http://<host>:<port>/v1` |\n| API key | none (local) |\n| Port | `<port>` |\n| Tensor parallel | `<tp>` |\n| Max context | `<context>` |\n| GPU | `<detected GPU>` |\n\nThen give a ready-to-run example using those exact values:\n\n```bash\ncurl -s http://<host>:<port>/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"<model_id>\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\n## Remote vs. local\n\nAll scripts accept `--host user@hostname`. When given, they SSH to the target.\nSet `ROCM_SSH_HOST` and `ROCM_SSH_USER` env vars to avoid passing `--host`\nevery time.\n\nFor remote Docker commands, run them over SSH:\n```bash\nssh user@host 'docker run -d ...'\n```\nUse `localhost` for health/warmup curl URLs (curl runs on the remote host).\n\n## Gotchas\n\n**`CUDA_VISIBLE_DEVICES` set to empty string** -- ROCm maps this variable to\n`HIP_VISIBLE_DEVICES`. Setting it to an empty string hides all GPUs.\n`CUDA_VISIBLE_DEVICES=0,1` works fine for restricting GPUs (same as\n`HIP_VISIBLE_DEVICES=0,1`). If the host has it set to empty, unset it:\n`unset CUDA_VISIBLE_DEVICES`. Do not pass `--env CUDA_VISIBLE_DEVICES=` (empty)\ninto Docker -- that also hides all GPUs inside the container.\n\n**FP4BMM crash on gfx942 (MI300X)** -- If the container exits immediately\nwith a segfault or illegal instruction: `VLLM_ROCM_USE_AITER_FP4BMM` must be\n`0` on gfx942. This is set correctly in `gpu_overrides.json` for gfx942.\nSee vLLM issue #34641.\n\n**`HIP error: no kernel image`** -- The Docker image has no compiled kernel\nfor your GPU's gfx version. Use `vllm/vllm-openai-rocm:latest`; it includes\ngfx942 and gfx950 kernels.\n\n**MLA models need `--block-size 1`** -- DeepSeek-R1/V3, Kimi-K2.5.\nWithout it the MLA attention backend silently falls back to a slower path.\nThis is in the recipe args for these models.\n\n**MoE models on multi-GPU need `--distributed-executor-backend mp`** --\nQwen3-235B, GLM-4.5, MiniMax-M2. The default distributed executor does not\nwork reliably with MoE on ROCm.\n\n**OOM during HIP graph capture** -- If the container logs show \"out of memory\"\nafter \"capturing CUDA graphs\" or \"capturing HIP graphs\", the model fits in\nVRAM but there isn't enough headroom for graph capture. Retry with\n`--enforce-eager` added to the vLLM args. This disables graph capture and\nfrees 1-2 GB. Trade-off: slightly higher decode latency, but the model runs.\n\n**\"Engine core initialization failed\"** -- This opaque error means the engine\ncore subprocess died. Check early container logs: `docker logs <name> 2>&1 |\nhead -50`. Common causes: gated model access denied (license not accepted on\nHF), unsupported architecture on this vLLM version, OOM during weight loading,\nmissing `--trust-remote-code` for custom architectures, or vLLM version too old\nfor the model (check `min_vllm_version` in the recipe).\n\n**`/dev/kfd` permission denied** -- User is not in the `video` or `render`\ngroup. Fix: `sudo usermod -aG video,render $USER` (requires re-login).\n\n**SSH key not configured** -- The scripts use `BatchMode=yes` SSH. If SSH\nfails with `Permission denied (publickey)`, configure key-based access first.\n\n**Restricting GPUs on shared hosts** -- Use `--env HIP_VISIBLE_DEVICES=0,1`\nor `--env CUDA_VISIBLE_DEVICES=0,1` to target specific GPUs by index.\n`HIP_VISIBLE_DEVICES` is the canonical AMD variable; `CUDA_VISIBLE_DEVICES`\nalso works (ROCm maps it). Never set either to an empty string.\n\n---\n\n## Reference\n\nPrecision compatibility, VRAM estimation, Docker flags, and known quirks:\n[reference.md](reference.md)\n"
}

SHA-256: b5436b0afdb95f73e36130e10aa4e9bad4573f9decc3d4151c93f84903dfd52b