{"id":17596,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.672Z","digest":"b157e8eb6f53ab6f3c81e941078c01b75f282a4abf6eea62c59cf926b7cbd281","against":null,"payload":{"name":"kermt-setup","description":"Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test inside the container. Every other kermt-* skill depends on this; invoke it first.","included_files":[],"skill_md_contents":"---\nname: kermt-setup\ndescription: Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test inside the container. Every other kermt-* skill depends on this; invoke it first.\nlicense: Apache-2.0\ncompatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.\nmetadata:\n  owner: evax@nvidia.com\n  classification: atomic-skill\n  risk_tier: skill\n# This file is intentionally short (~110 lines, ~1200 tokens) — well within the\n# 500-line / 5000-token budget for skill files. Longer reference material lives\n# alongside agent/scripts/kermt_container.sh.\n---\n\n# kermt-setup\n\nBootstrap the KERMT agent environment. Run this once on a fresh machine (or\nafter the Dockerfile or `environment.yml` changes) before invoking any other\n`kermt-*` skill.\n\n## Hardware requirements\n\n- **GPU**: at least one CUDA-capable NVIDIA GPU visible to the host. The image\n  is based on `nvidia/cuda:12.6.3-cudnn-devel-ubuntu22.04`, so the host driver\n  must support CUDA 12.6. Verify with host `nvidia-smi` before invoking.\n- **Host docker**: docker engine + nvidia-container-toolkit. Without the\n  toolkit, `docker run --gpus all` will fail at step 2 of the workflow below.\n- **Disk**: ≈ 50 GB free for the built kermt image (`docker image inspect\n  --format '{{.Size}}'` reports ≈ 44 GB; the `docker images` Size column\n  can show ~100 GB because it counts shareable buildx attestation layers\n  that are deduplicated across images). Plan for ~50 GB of unique on-disk\n  storage; add a comfortable buffer if you're also keeping build cache.\n- **Memory**: the build itself peaks at ~4 GB RAM during conda env solve.\n- This skill does not run training/inference workloads itself; per-workflow\n  hardware requirements (VRAM, GPU count) are declared in the respective\n  `kermt-<workflow>` skills.\n\n## When to invoke\n\n- User explicitly asks (`/kermt-setup`, \"set up kermt\", \"build the kermt image\",\n  etc.).\n- Or another `kermt-*` skill detected that the image does not exist and routed\n  here. (Most other skills call `kermt_ensure_image` themselves, so this is\n  usually only needed for the first-time setup, debugging, or a forced rebuild.)\n\n## Inputs\n\nThe skill takes no required arguments. Optional overrides (via env vars before\ninvoking, or by setting them in the user's shell):\n\n- `KERMT_IMAGE` — image tag to build/verify (default: `kermt:latest`).\n- `KERMT_REPO` — host path of the kermt repo checkout (default: auto-derived\n  from the script's location).\n\nIf the user has not specified a repo path and the current working directory is\nnot inside a kermt repo clone, ask for the repo path before proceeding.\n\n## Workflow\n\nAll work goes through `agent/scripts/kermt_container.sh`. The script's\nsubcommand dispatch can be invoked directly without sourcing — that is the\npreferred form for skill use.\n\nLet `HELPER=$KERMT_REPO/agent/scripts/kermt_container.sh`.\n\n1. **Verify docker is installed and the daemon is reachable.**\n   ```\n   $HELPER check_docker\n   ```\n   Exit 0 → continue. Non-zero → surface the error to the user (typically\n   \"docker not on PATH\" or \"daemon not reachable\"); do not attempt step 2.\n\n2. **Verify GPU passthrough works.**\n   ```\n   $HELPER check_gpu\n   ```\n   This runs `docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04\n   nvidia-smi` and checks the exit status. Non-zero → tell the user to install\n   `nvidia-container-toolkit` on the host and confirm a CUDA-capable NVIDIA GPU\n   is visible to the host (`nvidia-smi` on the host should also work). Stop\n   here; without GPU passthrough the kermt image will build but no workflow\n   will run.\n\n3. **Build or verify the kermt image.**\n   ```\n   $HELPER ensure_image\n   ```\n   If the image already exists, this returns immediately. Otherwise it builds\n   from `$KERMT_REPO/Dockerfile`. **Warn the user before invoking** that the\n   first build takes ~10–20 minutes on a typical workstation and streams build\n   logs to the console. Do not run this in the background — the user wants to\n   see progress and any build failures must surface immediately.\n\n4. **GPU smoke test inside the container.** Quote the whole `python` command\n   as a single string — the helper passes args through `bash -c \"$*\"`, so\n   unquoted multi-word commands get re-parsed and any embedded quotes are\n   collapsed.\n   ```\n   $HELPER run -- 'python -c \"import torch; print(\\\"cuda_available:\\\", torch.cuda.is_available()); print(\\\"device_count:\\\", torch.cuda.device_count())\"'\n   ```\n   Expected output: `cuda_available: True` and a positive `device_count`. If\n   `cuda_available` is `False` despite step 2 passing, something is wrong with\n   the container's CUDA wiring — report the full output to the user and stop;\n   do not declare the environment ready.\n\n5. **Summary to user.** Report:\n   - Image tag and ID (`docker image inspect $KERMT_IMAGE --format '{{.Id}}'`).\n   - Image size (`docker image inspect $KERMT_IMAGE --format '{{.Size}}'`).\n   - GPU count detected inside the container.\n   - \"Ready\" — the user can now invoke other `kermt-*` skills.\n\n## Hard rules\n\n- Do **not** pull or push docker images. The kermt image is built locally only.\n- Do **not** auto-delete or prune older `kermt:*` tags without the user's\n  explicit confirmation — the user may be running a finetune or pretrain in\n  another container that depends on a specific tag.\n- Do **not** modify the host's docker daemon configuration, daemon.json, or\n  user-group membership.\n- Do **not** modify the `Dockerfile` or `environment.yml` as part of this\n  skill. If the build fails because of a Dockerfile issue, surface the error\n  and stop; let the user decide whether to edit.\n- Do **not** rebuild the image when it already exists (i.e. do not pass a\n  `--no-cache` or `--pull` flag to ensure_image) unless the user explicitly\n  asks for a forced rebuild.\n\n## Forced rebuild\n\nIf the user explicitly asks to rebuild (e.g. after changing the Dockerfile or\n`environment.yml`), the cleanest path is to remove the old image first, then\nrerun `ensure_image`:\n\n```\ndocker image rm $KERMT_IMAGE\n$HELPER ensure_image\n```\n\nConfirm with the user before running `docker image rm`.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}