← Runpod (Official)CONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Runpod (Official)
Snapshot Sep 30, 2026 · 23:02 UTC · version 1.1.2
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "runpod",
"description": "Start here for any Runpod task — running GPU/CPU pods, deploying serverless endpoints, templates, network volumes, building images, or understanding how Runpod works. Routes the request to the right Runpod skill (runpod-mcp, runpodctl, flash, companion-clis, or runpod-usage). Use when it is unclear which Runpod skill applies.",
"included_files": [
{
"relative_path": "evals/mcp-vs-runpodctl-tiebreak.eval.md",
"size_in_bytes": 1312
},
{
"relative_path": "evals/prefer-prebuilt.eval.md",
"size_in_bytes": 1144
},
{
"relative_path": "evals/route-service-vs-endpoint.eval.md",
"size_in_bytes": 1154
},
{
"relative_path": "golden-paths/01-ollama-pod.md",
"size_in_bytes": 12969
},
{
"relative_path": "golden-paths/02-comfyui-pod/README.md",
"size_in_bytes": 5655
},
{
"relative_path": "golden-paths/02-comfyui-pod/variant-a-from-scratch.md",
"size_in_bytes": 5786
},
{
"relative_path": "golden-paths/02-comfyui-pod/variant-b-prebuilt.md",
"size_in_bytes": 6143
},
{
"relative_path": "golden-paths/03-whisper-endpoint/README.md",
"size_in_bytes": 6156
},
{
"relative_path": "golden-paths/03-whisper-endpoint/variant-a-hub.md",
"size_in_bytes": 5829
},
{
"relative_path": "golden-paths/03-whisper-endpoint/variant-b-flash.md",
"size_in_bytes": 5804
},
{
"relative_path": "golden-paths/04-finetune-pod.md",
"size_in_bytes": 13985
},
{
"relative_path": "golden-paths/05-model-to-endpoint-pipeline.md",
"size_in_bytes": 11430
},
{
"relative_path": "golden-paths/06-dev-pod.md",
"size_in_bytes": 12249
},
{
"relative_path": "golden-paths/07-network-volume-handoff.md",
"size_in_bytes": 9712
},
{
"relative_path": "golden-paths/08-finetune-to-serverless.md",
"size_in_bytes": 13413
},
{
"relative_path": "golden-paths/09-custom-serverless-dev-loop/README.md",
"size_in_bytes": 18270
},
{
"relative_path": "golden-paths/09-custom-serverless-dev-loop/template/Dockerfile",
"size_in_bytes": 1711
},
{
"relative_path": "golden-paths/09-custom-serverless-dev-loop/template/handler.py",
"size_in_bytes": 3848
},
{
"relative_path": "golden-paths/09-custom-serverless-dev-loop/template/requirements.txt",
"size_in_bytes": 500
},
{
"relative_path": "golden-paths/09-custom-serverless-dev-loop/template/start.sh",
"size_in_bytes": 1564
},
{
"relative_path": "golden-paths/10-multi-region-ha-serverless.md",
"size_in_bytes": 22260
},
{
"relative_path": "golden-paths/11-public-endpoints.md",
"size_in_bytes": 11007
},
{
"relative_path": "golden-paths/12-serverless-streaming.md",
"size_in_bytes": 8848
},
{
"relative_path": "golden-paths/13-autoscaling-tuning.md",
"size_in_bytes": 15731
},
{
"relative_path": "golden-paths/14-load-balancing-endpoint.md",
"size_in_bytes": 15347
},
{
"relative_path": "golden-paths/15-monitor-and-debug.md",
"size_in_bytes": 14689
},
{
"relative_path": "golden-paths/16-serverless-webhooks.md",
"size_in_bytes": 10015
},
{
"relative_path": "golden-paths/17-serverless-websocket.md",
"size_in_bytes": 19955
},
{
"relative_path": "golden-paths/18-concurrent-handler.md",
"size_in_bytes": 11767
},
{
"relative_path": "golden-paths/19-three-region-same-file.md",
"size_in_bytes": 17371
},
{
"relative_path": "golden-paths/20-model-caching-endpoint.md",
"size_in_bytes": 9443
},
{
"relative_path": "golden-paths/21-storage-tiers.md",
"size_in_bytes": 4689
},
{
"relative_path": "golden-paths/22-minimal-pod-image/README.md",
"size_in_bytes": 4979
},
{
"relative_path": "golden-paths/22-minimal-pod-image/template/Dockerfile",
"size_in_bytes": 803
},
{
"relative_path": "golden-paths/22-minimal-pod-image/template/start.sh",
"size_in_bytes": 805
},
{
"relative_path": "golden-paths/23-minimal-queue-image/README.md",
"size_in_bytes": 4952
},
{
"relative_path": "golden-paths/23-minimal-queue-image/template/Dockerfile",
"size_in_bytes": 436
},
{
"relative_path": "golden-paths/23-minimal-queue-image/template/handler.py",
"size_in_bytes": 321
},
{
"relative_path": "golden-paths/23-minimal-queue-image/template/requirements.txt",
"size_in_bytes": 7
},
{
"relative_path": "golden-paths/25-bake-vs-mount/README.md",
"size_in_bytes": 6446
},
{
"relative_path": "golden-paths/25-bake-vs-mount/template/Dockerfile",
"size_in_bytes": 796
},
{
"relative_path": "golden-paths/25-bake-vs-mount/template/start.sh",
"size_in_bytes": 426
},
{
"relative_path": "golden-paths/README.md",
"size_in_bytes": 11028
}
],
"skill_md_contents": "---\nname: runpod\ndescription: >-\n Start here for any Runpod task — running GPU/CPU pods, deploying serverless\n endpoints, templates, network volumes, building images, or understanding how\n Runpod works. Routes the request to the right Runpod skill (runpod-mcp,\n runpodctl, flash, companion-clis, or runpod-usage). Use when it is unclear\n which Runpod skill applies.\nmetadata:\n author: runpod\n version: \"1.1.2\" # x-release-please-version\nlicense: Apache-2.0\n---\n\n# Runpod (router)\n\nThe entrypoint for the Runpod skills. This skill does no work itself — it picks\nthe right lane and hands off. Read the matching skill's `SKILL.md` next.\n\n## The lanes\n\n| Lane | Use it for |\n| --- | --- |\n| **runpod-mcp** | Manage infra (pods, endpoints, jobs, templates, volumes, registries, catalog, billing) via **structured tool calls** — when the Runpod MCP tools are connected in this session. |\n| **runpodctl** | Manage the same infra from a **terminal/CI/script**, plus the things only the CLI does: Hub browse/deploy, `send`/`receive` file transfer, SSH keys, `doctor` setup, model cache. |\n| **flash** | **Write Python** that runs on Runpod serverless — `@remote`/`@Endpoint` functions, `flash dev` hot-reload, `flash deploy`. Code-first, not infra management. |\n| **companion-clis** | **Prerequisite artifacts**: download a model (`hf`), build/push an image (`docker`), repos/releases (`gh`), move data to a network volume over S3 (`aws`). |\n| **runpod-usage** | **Understand** how Runpod works before acting — pods vs serverless, building a container, storage, GPU selection, gotchas. Knowledge only. |\n\n## First run — check auth before the first infra action\n\nInfra tasks (pods, endpoints, jobs, volumes) need a working control plane — the **Runpod MCP**\nor **runpodctl**. Don't start and discover mid-task that nothing's set up: check first, and if\nit isn't, help the user set up rather than limping on a partial fallback.\n\n**Check** (credential resolution order: `RUNPOD_API_KEY` env → `.env` → `~/.runpod/config.toml`):\n```bash\nrunpodctl user # succeeds ⇒ a key is set and valid\n```\nPlus, in Claude Code, `/mcp` should show `runpod` **Connected**.\n\n**Rule: get a key first — do not default to MCP OAuth.** The reason: one `RUNPOD_API_KEY`\nunlocks every tool — it authenticates **runpodctl + flash + the hosted MCP** (as `--header\n\"Authorization: Bearer $RUNPOD_API_KEY\"`). The MCP's \"Sign in with Runpod\" OAuth auths the **MCP alone** — the CLIs\nstay blocked, so you hit a wall on any CLI-only task (Hub, `send`/`receive`, SSH, `doctor`,\nmodel cache/Model Repository, CPU endpoints). ⚠️ **OAuth-only is a half-setup.** If nothing's\nset up, stop and get a key, in order:\n1. **`flash login`** — browser OAuth that saves a real key to `~/.runpod/config.toml` (runpodctl\n + flash read it; reuse it for the MCP Bearer). One step, unlocks all. Human-only.\n2. **`export RUNPOD_API_KEY=…`** (https://console.runpod.io/user/settings) — same full unlock;\n best for headless agents.\n3. **MCP OAuth only** (`/mcp` → *Sign in*) — last resort, MCP-only work; CLIs stay unauthed.\n\n**Then:** if a lane already works, use it — but if *only* the MCP is OAuth'd, still get a key\nbefore any CLI-only step. Missing a CLI? `curl -sSL https://cli.runpod.net | bash` (runpodctl) ·\n`uv tool install runpod-flash` (flash) · `npx @runpod/mcp-server@latest add` (MCP). Full setup:\n[`runpod-usage/reference/getting-started.md`](../runpod-usage/reference/getting-started.md).\n\n## How to route\n\n1. **Conceptual question, or an unmade design choice** (serverless vs pod? which\n GPU? bake the model or mount a volume?) → read **runpod-usage** first, then\n continue with the answer.\n2. **Write/iterate/ship your own code on Runpod GPUs** → **flash**.\n3. **Produce an artifact** (download a model, build+push an image, create a repo\n release, sync data to a volume) → **companion-clis**.\n4. **Manage infrastructure** (create/list/update/delete pods, endpoints,\n templates, volumes; list GPUs/data centers; run a serverless job; billing):\n - Capability only the CLI has — **Hub, `send`/`receive`, SSH keys, `doctor`,\n model cache** → **runpodctl**.\n - Otherwise, if the Runpod **MCP tools are connected** in this session\n (`create-pod`, `list-endpoints`, … are available) → **runpod-mcp**.\n - Otherwise (shell-only agent, no MCP) → **runpodctl**.\n\n### runpod-mcp vs runpodctl (the overlap)\n\nBoth drive the same Runpod API, so they overlap on infra CRUD. Choose by\n**capability first, environment second**:\n\n- **MCP wins on convenience** for simple, structured operations — reads and basic\n CRUD — when its tools are connected (typed params, no shell quoting).\n- **runpodctl takes over when an operation needs a capability MCP lacks** — even\n if MCP is connected — and is the only option for a shell-only agent or when the\n user wants a reproducible command.\n\nCapability matrix (pick the preferred lane per operation):\n\n| Operation | Preferred lane | Why |\n| --- | --- | --- |\n| List/get anything; start/stop/restart/delete a pod; simple CRUD on endpoints, templates, volumes, registries; catalog; billing | **runpod-mcp** if connected, else runpodctl | Simple structured ops — MCP is typed and convenient |\n| Create a **simple** pod (one image + one GPU) | **runpod-mcp** if connected, else runpodctl | Both handle it |\n| Create a pod **from a template** or a **CPU** pod | **runpod-mcp** if connected, else runpodctl | MCP's create-pod takes `templateId` (v2-only) and `computeType: \"CPU\"` |\n| Create a pod with a **multi-GPU priority list**, or **template + CPU together** | **runpodctl** | MCP narrows to one GPU type, and rejects a template deploy for a CPU pod |\n| Deploy from the **Hub** | **runpod-mcp** if connected, else runpodctl | MCP has `list-hub-repos` + `deploy-hub-repo` |\n| **File transfer** (`send`/`receive`), **SSH** keys/info, **`doctor`** setup, **model** cache | **runpodctl** | MCP has no tool for these |\n| Invoke a serverless job (`run`/`runsync`/status/stream) | **runpod-mcp** if connected, else runpodctl | MCP has first-class job tools |\n\nRule of thumb: **default to MCP for the easy stuff, hand off to runpodctl the\nmoment an op needs a flag/feature MCP doesn't expose.**\n\n## Deploying a workload (the golden loop)\n\nFor any \"get <X> running on Runpod\" task, follow the **development loop** in\n`runpod-usage/reference/development-loop.md`: decide pod vs serverless → provision → set up\n(only if from-scratch) → verify → deliver → cost-guard + teardown. Two rules bind within it:\n\n- **Prefer a prebuilt template / Hub worker over building an image from scratch.**\n- **Before delivering, verify the workload with a real request from outside the pod/endpoint\n — a \"Running\"/\"ready\" status does not mean it is serving.**\n\nIt branches to two sub-loops:\n\n- **Service you open at a URL** (Ollama, ComfyUI, dev box) →\n [`runpod-usage/reference/pod-workflows.md`](../runpod-usage/reference/pod-workflows.md)\n (ports + env + volume at creation, SSH-exec install, bind `0.0.0.0`, poll the\n proxy URL). Execute in the runpodctl lane.\n- **Request/response API that scales to zero** (Whisper, inference) →\n [`runpod-usage/reference/endpoint-workflows.md`](../runpod-usage/reference/endpoint-workflows.md)\n (Hub worker vs flash vs custom image; invoke `/run`/`/runsync`; poll job status).\n\n## Worked examples (golden paths)\n\n**Two dozen** end-to-end scenarios (nearly all **live-verified**) live in\n[`./golden-paths/README.md`](./golden-paths/README.md) — the yardstick for \"can\nan agent finish the job\", with real commands + observed output to copy from. When a\ntask matches one, **open its golden path first** instead of re-deriving it:\n\n| Want to… | Golden path |\n| --- | --- |\n| Run a server (Ollama/ComfyUI) on a pod at a URL | [01](./golden-paths/01-ollama-pod.md), [02](./golden-paths/02-comfyui-pod/README.md) |\n| Deploy a serverless model endpoint (Hub / flash / custom image) | [03](./golden-paths/03-whisper-endpoint/README.md), [05](./golden-paths/05-model-to-endpoint-pipeline.md) |\n| Serve a HuggingFace model without baking it in or a volume (host-cached) | [20 — model caching (`--model-reference`)](./golden-paths/20-model-caching-endpoint.md) |\n| Call a ready hosted model (no deploy) | [11 — Public Endpoints](./golden-paths/11-public-endpoints.md) |\n| Fine-tune, then serve the result | [04](./golden-paths/04-finetune-pod.md), [08](./golden-paths/08-finetune-to-serverless.md) |\n| Interactive dev box (SSH / VS Code) | [06](./golden-paths/06-dev-pod.md) |\n| Move data pod → volume → serverless | [07](./golden-paths/07-network-volume-handoff.md) |\n| **Custom serverless when flash isn't enough** (dual-mode image dev loop) | [09](./golden-paths/09-custom-serverless-dev-loop/README.md) |\n| Build a minimal image for a target (pod vs serverless queue) | [22 (pod)](./golden-paths/22-minimal-pod-image/README.md), [23 (queue)](./golden-paths/23-minimal-queue-image/README.md); concepts in [building-images](../runpod-usage/reference/building-images.md) |\n| Decide what to bake into the image vs mount on a network volume | [25 — bake vs mount](./golden-paths/25-bake-vs-mount/README.md) |\n| **High availability / multi-region** serverless (multi-volume + data sync) | [10](./golden-paths/10-multi-region-ha-serverless.md), [19 (3-region)](./golden-paths/19-three-region-same-file.md) |\n| Stream output incrementally (`/stream`) | [12](./golden-paths/12-serverless-streaming.md) |\n| Tune autoscaling / raise per-worker throughput | [13 (autoscaling)](./golden-paths/13-autoscaling-tuning.md), [18 (concurrency)](./golden-paths/18-concurrent-handler.md) |\n| Load-balancing / HTTP-server or WebSocket worker | [14 (LB)](./golden-paths/14-load-balancing-endpoint.md), [17 (WebSocket)](./golden-paths/17-serverless-websocket.md) |\n| Get notified on job completion (push, not poll) | [16 — webhooks](./golden-paths/16-serverless-webhooks.md) |\n| Check health / debug a failing endpoint | [15 — monitor & debug](./golden-paths/15-monitor-and-debug.md) |\n\n## Multi-lane tasks\n\nSequence is always **understand → produce artifacts → manage infra → verify**,\nbecause infra can only reference artifacts that already exist. Keep each step in\none lane, and switch lanes at credential boundaries.\n\nExample — \"deploy `openai/gpt-oss-20b` to a serverless endpoint\":\n1. **runpod-usage** — serverless vs pod, GPU tier for 20B, bake vs mount vs cache.\n2. **companion-clis** — `hf download …`, `docker build --platform=linux/amd64 …`, `docker push`.\n3. **runpod-mcp** or **runpodctl** — create the endpoint referencing the image + GPU pool.\n4. Same infra lane — invoke the endpoint / check status to verify.\n\n## Auth\n\nEverything is one key: **`RUNPOD_API_KEY`** (https://console.runpod.io/user/settings).\nEach lane just makes that key resolvable — `runpodctl doctor`, `flash login`, MCP\nstdio env var, or MCP hosted \"Sign in with Runpod\" (OAuth, no key on disk).\nCompanion CLIs use their **own** credentials (HuggingFace token, GitHub auth,\nDocker Hub PAT, Runpod **S3** keys for `aws`) — do not reuse `RUNPOD_API_KEY` for\nthose.\n"
}SHA-256: 8f086b9c4f9da2356faedf01b0dc14cd52b5970acf093a7556a133bae3100f45