← AMDCONTENT HISTORY

Update to AMD

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.2.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "hyperloom-workload-optimizer",
  "description": "Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. Given a model, framework, workload (TP/EP, concurrency, ISL/OSL, precision), an objective and a time budget, it explores per-workload which levers to pull (serving/config parameters and env, framework enablement and source patches, and hot GPU-kernel rewrites), benchmarks each candidate, and returns the optimization stack that produced the gain. Use when the user wants to make a model serve faster, raise tokens/sec or throughput, optimize or tune vLLM or SGLang on MI300X/MI325X/MI355X, run Hyperloom, run the kernel-agent, quantize-then-optimize with Quark, set up Hyperloom from scratch, or resume a Hyperloom session. Do not use to stand up a server for plain serving, diagnose a broken ROCm install, or run a one-off kernel/benchmark or trace analysis without the optimization loop.",
  "included_files": [
    {
      "relative_path": "evals/evals.json",
      "size_in_bytes": 5089
    },
    {
      "relative_path": "evals/machine.yml",
      "size_in_bytes": 22
    },
    {
      "relative_path": "reference.md",
      "size_in_bytes": 5328
    },
    {
      "relative_path": "scripts/_env.sh",
      "size_in_bytes": 2349
    },
    {
      "relative_path": "scripts/launch.sh",
      "size_in_bytes": 1846
    },
    {
      "relative_path": "scripts/launch_health.sh",
      "size_in_bytes": 2052
    },
    {
      "relative_path": "scripts/preflight.py",
      "size_in_bytes": 9826
    },
    {
      "relative_path": "scripts/resume.sh",
      "size_in_bytes": 1862
    },
    {
      "relative_path": "scripts/tests/test_launch_flow.sh",
      "size_in_bytes": 9045
    },
    {
      "relative_path": "scripts/tests/test_preflight.py",
      "size_in_bytes": 6609
    },
    {
      "relative_path": "setup.md",
      "size_in_bytes": 5504
    },
    {
      "relative_path": "skill-card.md",
      "size_in_bytes": 203
    }
  ],
  "skill_md_contents": "---\nname: hyperloom-workload-optimizer\ndescription: >-\n  Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs\n  and reports a validated gain, using the Hyperloom multi-agent optimizer. Given a\n  model, framework, workload (TP/EP, concurrency, ISL/OSL, precision), an objective\n  and a time budget, it explores per-workload which levers to pull (serving/config\n  parameters and env, framework enablement and source patches, and hot GPU-kernel\n  rewrites), benchmarks each candidate, and returns the optimization stack that\n  produced the gain. Use when the user wants to make a model serve faster, raise\n  tokens/sec or throughput, optimize or tune vLLM or SGLang on MI300X/MI325X/MI355X,\n  run Hyperloom, run the kernel-agent, quantize-then-optimize with Quark, set up\n  Hyperloom from scratch, or resume a Hyperloom session. Do not use to stand up a\n  server for plain serving, diagnose a broken ROCm install, or run a one-off\n  kernel/benchmark or trace analysis without the optimization loop.\n---\n\n<!--\nCopyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.\n\nSee LICENSE for license information.\n-->\n\n# Hyperloom Workload Optimizer\n\nYou are the catalog entry point for Hyperloom optimization on AMD Instinct GPUs.\nBootstrap the workspace, prepare the runtime environment, collect workload\nparameters, then install, launch, and monitor the optimizer. This skill owns the\norchestration and the launcher gates; environment prep and workload intake are\ndelegated to the skills the Hyperloom wheel installs, and\n`@${HYPERLOOM_SKILL_PATH}` (`inference_optimizer`) is the execution baseline.\n\nDo not manually optimize inside chat unless debugging.\n\n## Prerequisites\n\n- AMD Instinct GPU host (MI300X / MI325X / MI355X) with ROCm\n- `/dev/kfd` and `/dev/dri` present; `amd-smi` or `rocm-smi` works\n- Python 3.10+ and network access to install the Hyperloom wheel\n- Anthropic (or compatible) LLM credentials for agent backends\n- A dedicated agent workspace directory\n\nEvery command in this skill runs on that GPU host. Confirm the shell you are in\nis on it before Phase 0, so a bootstrap does not land on a machine with no GPU.\n\nThe Hyperloom **runtime** ships via `pip install` of the published wheel.\n\n## What Hyperloom runs\n\nThe CLI starts a Python Coordinator that coordinates:\n\n- **Orchestration** — baseline, explore, specialist, integrate_patch, sweep\n- **Kernel** — trace_analyze, run_optimization, integrate\n- **Critic** — proposal review (default `--critic-agent`)\n- **Robustness** — health monitoring and RCA (default `--robustness-agent`)\n\nState lives under a **session directory** per run; run-state root is\n`$USER_DATA_PATH` (default `/workspace/hyperloom`), independent of the install\ndirectory (`INSTALL_DIR`, where the wheel and `.env` live) and may point to\nshared storage. Layout: `$USER_DATA_PATH/runtime/` (install.sh outputs,\n`kernel-agent.env.sh`), `logs/`, and `<model_basename>/<UTC_ts>/` per session\nholding `manifest.json`, `state.json`, `runs/`, `reports/`, `optimizer_runs/`.\n\n## Workflow overview\n\nMatch `hyperloom-custom-advanced` section order — do **not** ask workload\nquestions while writing `.env` or during `/hyperloom-setup`.\n\n- **Phase 0 Bootstrap** — `pip install`, `/hyperloom-setup` → `.env` (credentials + run mode only)\n- **Phase 1 Environment** — custom-advanced §Setup Configuration (baremetal: confirm host; docker: start container + setup inside, contract in [setup.md](setup.md))\n- **Phase 2 Workload intake** — custom-advanced §Advanced Configuration → Model Resolution → show launch plan → user confirms\n- **Phase 3 Execute** — install.sh → preflight → launch → monitor → report\n\nLoad `hyperloom-custom-advanced` at Phase 1 and follow its sections in order\n(discovery: `.cursor/` / `.claude/` / `.agents/skills/hyperloom-custom-advanced/SKILL.md`).\nIf it is not on disk, stop and tell the user to restart the agent so the newly\ninstalled skills are picked up — do not improvise the environment or workload\nsections from memory, since the wheel is the source of truth for both.\nFor deeper optimizer behavior read `@${HYPERLOOM_SKILL_PATH}` (`inference_optimizer`);\nIron Rules + CLI reference: [reference.md](reference.md).\n\n## Iron Rules (launcher gates)\n\nRun order is always **IR-2 → IR-1 → launch**. Full text in [reference.md](reference.md).\n\n- **IR-1 — GPU unoccupied.** Before every `optimize` (fresh or `--resume`), every\n  visible GPU must have zero foreign serving PIDs (`sglang.launch_server` /\n  `vllm.entrypoints` / `Magpie`) and ≲ 500 MiB VRAM in use.\n- **IR-2 — install.sh before launch.** Run `install.sh` and source\n  `kernel-agent.env.sh` in the **same shell** that spawns `optimize`.\n- **Resume carve-out:** `--resume` may skip install only when `install.sh` exited\n  0 earlier in the same shell, `kernel-agent.env.sh` is still sourced, and the\n  session's `manifest.json` exists. Any failure → re-run `install.sh`.\n\n## Phase discipline (do not skip)\n\nOne phase at a time. Each phase asks only its own questions, waits for the\nuser's answers, completes its exit condition, then moves on. Never batch\nquestions from different phases into one prompt. In particular, never ask\nworkload questions (model, framework, TP/EP, precision, ISL/OSL, hours…) during\nPhase 0 or Phase 1 — those belong to Phase 2 only.\n\n## Phase 0 — Bootstrap\n\nSkip completed steps (idempotent). Ask only about the install directory and\ncredentials/run mode here. Do not ask about the model or workload yet.\n\n### Confirm the install directory\n\nThe wheel installs into a target directory with `pip install --target <dir>`,\nwhich also holds `.env` and runtime artifacts. Do not silently use the current\ndirectory. Show the resolved current directory (`pwd`) and confirm it with the\nuser, or let them choose another dedicated path. Wait for the answer, then `cd`\ninto the chosen directory before installing.\n\n### Install the Hyperloom wheel\n\nSkip when `hyperloom/` (wheel) or `src/hyperloom/` (source) already exists in the\nconfirmed directory.\n\nThe runtime is published to PyPI as `hyperloom-inference-optimizer`. List the\nreleases, tell the user the newest one, and ask whether to install it or a\nversion they name.\n\nList with `--pre` so prereleases are visible, and install an exact `==` version\nso a later bootstrap installs the same runtime.\n\n```bash\ncd \"$INSTALL_DIR\"   # the directory confirmed above\npip index versions hyperloom-inference-optimizer --pre\npip install hyperloom-inference-optimizer==<version the user approved> --target .\n```\n\nConfirm `hyperloom/inference_optimizer/assets/install.sh` exists. Restart the\nagent if wheel skills are not visible.\n\n### Credentials and run mode\n\nRun `/hyperloom-setup` (installed to `.cursor/skills/hyperloom-setup/`). It\nwrites `.env`, sets `USER_DATA_PATH`, `HYPERLOOM_RUN_MODE`, and\n`HYPERLOOM_SKILL_PATH`, and on bare metal runs `install_baremetal.sh`.\n\n**Phase 0 is done when all hold:**\n\n- `hyperloom/inference_optimizer/assets/install.sh` exists\n- `.env` exists with non-placeholder LLM secrets\n- `USER_DATA_PATH`, `HYPERLOOM_RUN_MODE`, and `HYPERLOOM_SKILL_PATH` are set\n\nMore bootstrap detail: [setup.md](setup.md).\n\n## Phase 1 — Environment prep\n\nLoad `hyperloom-custom-advanced` and follow its **Setup Configuration**\nsection only.\n\n**Baremetal (`HYPERLOOM_RUN_MODE=baremetal`):** confirm `install_baremetal.sh`\nfinished and the serving framework from setup is importable. Do not ask workload\nquestions yet.\n\n**Docker (`HYPERLOOM_RUN_MODE=docker`):** image choice, `docker run`, and the\nin-container setup are owned entirely by custom-advanced Setup Configuration —\nfollow it, do not restate its commands or flags here. Do not ask workload\nquestions until the container is up and in-container setup succeeded, and never\nrun `optimize` on the host.\n\nPhase 1 is done when the target environment (host or container) is ready.\n\n## Phase 2 — Workload intake\n\nEnter only after Phase 0 and Phase 1 exit conditions hold. This is the first and\nonly phase that asks workload questions.\n\nNow follow custom-advanced **Advanced Configuration**, **Default Values**, and\n**Model Resolution**. Use the agent's structured question UI when available.\nNever copy API keys into chat output.\n\n| Field | CLI flag | Default | Notes |\n|---|---|---|---|\n| Model path | `--model` | required | Local dir with `config.json`, or HF cache |\n| Framework | `--framework` | `sglang` | or `vllm`; prefer `.env` `FRAMEWORK` when set |\n| TP / EP | `--tp` / `--ep` | `1` / `1` | tensor / expert parallel |\n| CONC | `--conc` | `64` | client concurrency |\n| ISL / OSL | `--isl` / `--osl` | `1024` / `1024` | input / output seq lengths |\n| PRECISION | `--precision` | `bf16` | match checkpoint; `fp8` for FP8 models |\n| MAX_HOURS | `--max-hours` | CLI `2.0` | offer `3` (quick) or `12` (full); see below |\n| TARGET_GAIN | `--target-gain` | `30` | desired % gain |\n\n**Optional:** `--no-explore`, `--no-enable-conc-sweep`, `--gpu-type`,\n`--server-args`, `--compare-against-gpu`, `--quantize` prelude.\n\nInfer `PRECISION` from the model name when obvious (e.g. an `FP8` model implies\n`--precision fp8`) and confirm it — do not silently keep the `bf16` default.\n\n### Budget and flags — offer these three\n\nOffer all three and let the user pick one. The flags in each are a set: pass them\ntogether, and do not ask for a budget and then ask separately which phases to run.\nThe two demos take the workload and flags of the Hyperloom demo skill of the same\nbudget — treat those as given and skip the table above. The user may name their\nown model instead of the demo's; for the 3-hour demo keep it at 8B or below.\nConfirm everything in the launch plan. Only **Custom** collects workload answers.\n\n**1. 3-hour demo** (`hyperloom-qwen3-8b-3h`) — `Qwen/Qwen3-8B` unless the user\nnames another 8B-or-smaller model, TP=1, CONC=64, ISL=OSL=1024,\n`--precision bf16`, serving and config parameters only, no kernel rewrites.\nResolve the model per custom-advanced Model Resolution; download it from Hugging\nFace when it is not already local. Match `--precision` to the chosen checkpoint.\nExpect a modest validated gain, or an honest 0% when the workload has no\nparameter headroom.\n\n```text\n--max-hours 3 --precision bf16\n--no-framework-agent --no-kernel --no-enable-conc-sweep --no-enable-roofline\n--max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01\n--explore-force-exit-budget-pct 0.01 --explore-force-exit-hours-remaining 0.05\n```\n\n**2. 12-hour demo** (`hyperloom-qwen3-14b-fp8-12h`) — `Qwen/Qwen3-14B-FP8`\nunless the user names another model, TP=1, CONC=64, ISL=OSL=1024,\n`--precision fp8` matched to the chosen checkpoint, every lever with kernel\nrewrites included. The kernel agent needs room to profile, rewrite and\nrevalidate, which is where the larger gains come from.\n\n```text\n--max-hours 12 --precision fp8\n--max-minutes-framework-pct 0.01 --max-minutes-explore-pct 0.42\n--max-minutes-kernel-pct 0.42\n```\n\n**3. Custom** — the user brings their own model or workload instead of taking a\ndemo. Walk through the fields in the table above and the phase toggles, one\nquestion at a time, and derive the flags from the answers rather than asking for\nflags. Whichever levers they pick, a budget of 3 hours or less keeps the 3-hour\ndemo's flag set.\nOptional flags come from the list above; show the full flag list in the launch\nplan either way.\n\n### Confirmation gate (required before Phase 3)\n\nThe Coordinator has no in-loop `setup` / `classify` — a value not asked here is\nsilently lost to its default. Before running any Phase 3 command, present the\nfull launch plan (including defaulted fields) and get explicit user confirmation.\n\nPrint the plan in the reply body as this aligned block:\n\n```text\nLaunch plan — please confirm:\n  MODEL_PATH    /wekafs/models/Qwen3-14B-FP8\n  FRAMEWORK     vllm\n  TP=1  EP=1  CONC=64\n  ISL=1024  OSL=1024\n  PRECISION=fp8\n  MAX_HOURS=3     TARGET_GAIN=20%\n  profile       3-hour demo — no kernel, no framework agent, no roofline\n  flags         --no-framework-agent --no-kernel --no-enable-conc-sweep\n                --no-enable-roofline\n                --max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01\n                --explore-force-exit-budget-pct 0.01\n                --explore-force-exit-hours-remaining 0.05\n  RUN_MODE      baremetal\n```\n\nNever put the plan inside the confirmation prompt itself. A prompt renders as one\nwrapped paragraph, which collapses the alignment above into an unreadable blob the\nuser has to search for `MAX_HOURS` in. Keep the prompt to a single short question\nsuch as `Approve this launch plan?`, and if you offer a \"change something\" option,\nname the field to change rather than making the user retype it as free text.\n\nDo **not** run `install.sh` or launch `optimize` until the user approves this plan.\n\n### Persist the plan (required — shells do not share exports)\n\nAgent shells do not persist exports between calls, so write the confirmed values\nto `$RUN_DIR/workload.env` right after approval. Every Phase 3 block sources it;\nwithout this, launch silently falls back to `${TP:-1}` / `${CONC:-64}` defaults\nand `--model \"\"`. Fill each value from the approved plan.\n\n```bash\nexport USER_DATA_PATH=\"${USER_DATA_PATH:?run /hyperloom-setup first}\"\nexport RUN_DIR=\"${USER_DATA_PATH}/optimizer_runs\"\nmkdir -p \"$RUN_DIR\"\n# Quoted heredoc (<<'EOF'): values are written literally, so a MODEL_PATH with\n# spaces, $, or $(...) is not expanded or executed. Edit each value to the plan.\ncat > \"$RUN_DIR/workload.env\" <<'EOF'\nexport MODEL_PATH=/wekafs/models/Qwen3-14B-FP8\nexport FRAMEWORK=vllm\nexport TP=1\nexport EP=1\nexport CONC=64\nexport ISL=1024\nexport OSL=1024\nexport PRECISION=fp8\nexport MAX_HOURS=3\nexport TARGET_GAIN=20\n# The whole flag set for the approved profile, space-separated. The 3-hour\n# demo is shown; a 12-hour run swaps in its own set.\nexport OPT_FLAGS=\"--no-framework-agent --no-kernel --no-enable-conc-sweep --no-enable-roofline --max-minutes-explore-pct 0.39 --max-minutes-sweep-pct 0.01 --explore-force-exit-budget-pct 0.01 --explore-force-exit-hours-remaining 0.05\"\nEOF\n```\n\n## Phase 3 — Install (IR-2)\n\n`INSTALL_DIR` is the directory confirmed in Phase 0, the one holding `hyperloom/`\nand `.env`. Every Phase 3 block below rebuilds it from the current directory, so\nrun them from there; the check refuses a directory that is not it.\n\nResolve paths for wheel or source layout:\n\n```bash\nexport INSTALL_DIR=\"$(pwd -P)\"\n[ -d \"${INSTALL_DIR}/hyperloom\" ] || [ -d \"${INSTALL_DIR}/src/hyperloom\" ] || {\n  echo \"ERROR: ${INSTALL_DIR} holds no hyperloom/ -- cd to the Phase 0 install directory\" >&2; exit 1; }\nset -a; . \"${INSTALL_DIR}/.env\"; set +a\nexport USER_DATA_PATH=\"${USER_DATA_PATH:?USER_DATA_PATH missing}\"\n. \"${USER_DATA_PATH}/optimizer_runs/workload.env\"   # confirmed Phase 2 values\nexport PYTHONPATH=\"${INSTALL_DIR}:${PYTHONPATH:-}\"\nulimit -Sn 65536 || true\n\nINSTALL_SH=\"${INSTALL_DIR}/hyperloom/inference_optimizer/assets/install.sh\"\n[ -f \"$INSTALL_SH\" ] || INSTALL_SH=\"${INSTALL_DIR}/src/hyperloom/inference_optimizer/assets/install.sh\"\n\nbash \"$INSTALL_SH\"\n. \"${KERNEL_AGENT_ENV:-${USER_DATA_PATH}/runtime/kernel-agent.env.sh}\"\nexport PYTHONPATH=\"${INSTALL_DIR}:${PYTHONPATH:-}\"\n```\n\nIn Docker mode, run this inside the container.\n\n## Phase 3 — Preflight (IR-1)\n\n`install.sh` exports `$PYTHON`; the fallback below covers agent sandboxes that do\nnot persist exports between shell calls.\n\n```bash\nexport SKILL_DIR=\"${SKILL_DIR:?absolute path of the directory holding this SKILL.md}\"\n. \"${USER_DATA_PATH}/optimizer_runs/workload.env\"   # confirmed Phase 2 values\nexport PYTHON=\"${PYTHON:-$(command -v python3)}\"\n\"$PYTHON\" \"${SKILL_DIR}/scripts/preflight.py\"\n```\n\nThe gate exits non-zero — do not launch — when `MODEL_PATH` is missing or has no\n`config.json`, torch sees no GPU, a foreign serving process still holds a card,\nor any GPU holds more than `IR1_VRAM_LIMIT_MIB` (default 500) MiB.\n\nIt also blocks when VRAM cannot be read at all: no `amd-smi`/`rocm-smi` on\n`PATH`, a probe that exits non-zero, or output it cannot parse. An unreadable\nprobe cannot rule out a busy GPU, and a foreign process holding VRAM under a\ndifferent name would slip through. Confirm the GPUs are idle by hand before\nre-running with `IR1_ALLOW_UNVERIFIED_VRAM=1`.\n\nNever print API keys or tokens. `scripts/tests/test_preflight.py` covers the\nprobe shapes this gate must reject.\n\n## Phase 3 — Launch\n\nAfter IR-2 and IR-1 pass, launch. `setsid nohup` is required for runs longer than\n5 minutes, so the run outlives the agent shell.\n\n```bash\nexport INSTALL_DIR=\"$(pwd -P)\"\nexport SKILL_DIR=\"${SKILL_DIR:?absolute path of the directory holding this SKILL.md}\"\nbash \"${SKILL_DIR}/scripts/launch.sh\"\n```\n\nEvery workload value comes from the confirmed `workload.env`; the script has no\n`${VAR:-default}` fallbacks, so a missing value fails loudly instead of launching\na different config. Put any optional Phase 2 flags (`--no-kernel`, `--no-explore`,\n`--gpu-type`, `--model-class`, `--server-args`, `--compare-against-gpu`,\n`--quantize`, phase budget flags) into `OPT_FLAGS` in `workload.env`. `OPT_FLAGS`\nis word-split, so quote any flag value that contains spaces, e.g.\n`export OPT_FLAGS='--server-args \"--foo bar\"'`.\n\n### Launch health check (30 s after start)\n\nRequired after every launch and resume. The PID recorded at launch is the\n**setsid wrapper**, which exits immediately — it is NOT the optimizer. This reads\nthe real `.pid` and `.session_dir` from the launch-info JSON, rewrites the PID\nfile so the monitor watches the right process, and records both in\n`$RUN_DIR/last_launch.env` for the later phases.\n\n```bash\nexport INSTALL_DIR=\"$(pwd -P)\"\nexport SKILL_DIR=\"${SKILL_DIR:?absolute path of the directory holding this SKILL.md}\"\nbash \"${SKILL_DIR}/scripts/launch_health.sh\"\n```\n\nIt exits non-zero when the launch-info JSON never appeared, no optimizer process\ncan be found, or `session_dir` is still unset — inspect the reported run log in\nthose cases. Never guess `session_dir` from a timestamp; concurrent sessions\nshare `USER_DATA_PATH`.\n\n## Phase 3 — Monitor\n\nPoll at most every 5 minutes unless debugging a startup failure. Use the state\nreader the wheel ships rather than parsing `state.json` by hand — it also prints\nthe recent lifecycle events.\n\n```bash\nexport INSTALL_DIR=\"$(pwd -P)\"\n. \"${USER_DATA_PATH}/optimizer_runs/last_launch.env\"   # SESSION_DIR from launch\nSTATE_TOOL=\"${INSTALL_DIR}/hyperloom/inference_optimizer/tools/read_optimizer_state.py\"\n[ -f \"$STATE_TOOL\" ] || STATE_TOOL=\"${INSTALL_DIR}/src/hyperloom/inference_optimizer/tools/read_optimizer_state.py\"\n\"${PYTHON:-python3}\" \"$STATE_TOOL\" \"$SESSION_DIR\"\n```\n\nFor recent action counts grouped by category, the wheel also ships\n`tools/event_counts.py`, invoked the same way.\n\nReport session id + log path, `baseline_tput` / `current_best` /\n`cumulative_gain`, explore accepted/rejected, last kernel opt (correctness,\nspeedup, KEEP/REVERT), and process-alive vs `stop_reason`. See\n[reference.md](reference.md) Report fields.\n\n## Resume\n\nResume runs in a fresh shell. Re-run the IR-2 and IR-1 gates first, exactly as for\na fresh launch — the script does not re-check them.\n\n```bash\nexport INSTALL_DIR=\"$(pwd -P)\"\nexport SKILL_DIR=\"${SKILL_DIR:?absolute path of the directory holding this SKILL.md}\"\nbash \"${SKILL_DIR}/scripts/resume.sh\"\nbash \"${SKILL_DIR}/scripts/launch_health.sh\"\n```\n\nIt resumes the session recorded in `last_launch.env` and always passes\n`--resume-from` explicitly, because a bare `--resume` auto-picks the newest\nsession and can target the wrong run. Resume writes its own log\n(`run_resume-*.log`) so the original run log is preserved. Reuse the IR-2\ncarve-out rules; re-run `install.sh` if the shell or env changed.\n\n| `stop_reason` | Action |\n|---|---|\n| `time_exhausted` | `--resume` same session |\n| `no_more_leverage` | stop; resume only if user changes strategy |\n| `policy_loop` | inspect `policy_denial_history`; clear stale prunes |\n\n## Expected optimizer flow\n\n1. Establish `baseline_tput`.\n2. Coordinator runs roofline/profile analysis after baseline.\n3. `explore` tests serving parameters incrementally.\n4. Kernel-agent runs on hot paths with compile + correctness evidence.\n5. `sweep` validates concurrency around the best candidate.\n6. Final report under `$SESSION_DIR/reports/`.\n\n## When to defer\n\n- **Plain serving only** — use `serving-llms-on-instinct`.\n- **ROCm driver broken** — diagnose the ROCm stack first (e.g. a `rocm-doctor`\n  skill if published); do not start the optimizer on a broken driver.\n- **Edge cases** — read `@${HYPERLOOM_SKILL_PATH}` for multi-node, atom\n  framework (IR-8), critic/robustness backends, cache topology, and the\n  full failure matrix.\n\n## Further reading\n\n- Bootstrap detail: [setup.md](setup.md)\n- Iron Rules + CLI reference: [reference.md](reference.md)\n- Authoritative runtime skill: `hyperloom/inference_optimizer/SKILL.md`\n"
}

SHA-256: fd41c949ca26544a77e33a6b44c62497fcdbd94c5842e8d2fdf9ff52d5bcecd4