← Plugin catalog
Developer Tools

Runpod (Official)

Runpod v1.1.2

Publisher description

From the marketplace listing

Official Runpod agent skills and a bundled MCP server: a router that dispatches to runpodctl, flash, and companion CLIs, plus conceptual usage guidance and worked golden paths.

Language: English · Automatically detected from descriptions.

Files & skills

File archives

Plugin package256 files · 497 KBBrowse files →
Skill instructions
companion-clis4.03 KB

View saved version →

---
name: companion-clis
description: Companion CLIs for Runpod workflows — HuggingFace, GitHub, Docker, and AWS.
allowed-tools: Bash(hf:*), Bash(gh:*), Bash(docker:*), Bash(aws:*), Bash(ssh-keygen:*), Bash(ssh-add:*), Bash(ssh-agent:*)
compatibility: Linux, macOS, Windows
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Companion CLIs

Four CLIs commonly needed alongside Runpod. Each has its own **credentials + command reference** in [`reference/`](reference/) — plus a one-time `<cli>-setup.md` for install (only opened if the CLI isn't installed). Load only the one the task needs, not all four.

| CLI | Use it to | Full reference |
|-----|-----------|----------------|
| `hf` (HuggingFace) | Download models from the Hub to cache/bake into images | [reference/huggingface.md](reference/huggingface.md) |
| `gh` (GitHub) | Manage worker repos + cut releases (Hub indexes releases) | [reference/github.md](reference/github.md) |
| `docker` | Build/validate/push images to Docker Hub for Runpod to pull | [reference/docker.md](reference/docker.md) |
| `aws` (S3) | Read/write network-volume storage over Runpod's S3 API | [reference/aws.md](reference/aws.md) |

Each requires credentials before use. Read the per-tool reference for auth steps and commands; install is a separate one-time `<cli>-setup.md`.

## Windows: Install WSL2 First

If you are on Windows, install WSL2 before proceeding — it gives you the native Linux environment all these CLIs target. In PowerShell as Administrator, then restart:

```powershell
wsl --install
```

Afterward open the Ubuntu app to finish setup, then follow the **Linux** instructions in each reference.

## HuggingFace CLI

Download models locally so they're cached for a Docker build/run. Auth and `hf download` recipes: **[reference/huggingface.md](reference/huggingface.md)** (install: [reference/huggingface-setup.md](reference/huggingface-setup.md)).

- Use the standalone `hf` CLI, **not** `pip install huggingface_hub` (that's the older `huggingface-cli` with different syntax).
- Auth via `hf auth login`, or `export HF_TOKEN=hf_...` (env var wins over saved token).

## GitHub CLI

Manage worker repositories and cut releases. Auth and commands: **[reference/github.md](reference/github.md)** (install + SSH-key setup: [reference/github-setup.md](reference/github-setup.md)).

- **The Hub indexes releases, not commits** — every Hub listing update needs a new `gh release create`.
- One SSH key (`ssh-keygen -t ed25519`) registers with both GitHub (`gh ssh-key add`) and HuggingFace (paste in browser).

## Docker

Build, validate, and push images to Docker Hub. Credentials and commands: **[reference/docker.md](reference/docker.md)** (install: [reference/docker-setup.md](reference/docker-setup.md)).

- **Always build `--platform=linux/amd64`** — Runpod runs on x86 Linux.
- **Always use explicit semantic tags; never `latest`** — `latest` doesn't track the newest push, so workers can silently pull the wrong image.
- Docker Hub auth uses a **personal access token**, not your password. For private images, register the credential once in Console → Container Registry Settings.

## AWS CLI

Access network-volume storage over Runpod's S3-compatible API (bucket name = network volume ID). Credentials, region rules, and commands: **[reference/aws.md](reference/aws.md)** (install: [reference/aws-setup.md](reference/aws-setup.md)).

- Runpod's S3 API, **not AWS**: access key = Runpod **user id** (`user_...`), secret = S3 API key (`rps_...`).
- **S3 API keys are Console-only.** No `runpodctl`/REST/GraphQL creates them — if they're not already in `~/.aws/credentials`/env and S3 access is needed, **stop and ask the user** to generate them (Settings > S3 API Keys).
- Every command needs `--region DATACENTER --endpoint-url https://s3api-DATACENTER.runpod.io/` (datacenter = the volume's DC, not an AWS region).
- For large/many-file transfers with reliable resume, see [reference/aws.md → optional resumable volume transfers](reference/aws.md#optional-resumable-volume-transfers-community-tool).

Referenced files: 8

flash12.1 KB

View saved version →

---
name: flash
description: >-
  runpod-flash — code-first serverless: write Python locally, run it on remote
  Runpod GPUs/CPUs with `flash dev` (hot-reload + live worker logs), then
  `flash deploy`. Use for @Endpoint/@remote functions, resource config, and
  debugging flash deployments. For CLI-only infra management use runpodctl or
  runpod-mcp.
user-invocable: true
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Runpod Flash

Write code locally, iterate with `flash dev` — it runs your functions on remote Runpod GPUs/CPUs with hot-reload and live worker logs — then `flash deploy` to ship. `Endpoint` handles provisioning.

**Load on demand — this skill keeps the mental model + gotchas inline; details live in [`reference/`](reference/):**

| Need | Read |
|------|------|
| Install, auth, `flash init`, and the full `flash` command list | [reference/setup-and-cli.md](reference/setup-and-cli.md) |
| `Endpoint(...)` constructor params, `NetworkVolume`/`PodTemplate`/`EndpointJob`, GPU & CPU enum tables | [reference/api.md](reference/api.md) |
| Worked patterns — choosing a model, warm-worker model loading, CPU→GPU pipeline, parallel calls | [reference/patterns.md](reference/patterns.md) |

Quick start: `uv tool install runpod-flash` → `flash login` (or `export RUNPOD_API_KEY=...`) → `flash init my-project` → `flash dev`. Details in [reference/setup-and-cli.md](reference/setup-and-cli.md).

## Dev vs Deploy

- `flash dev` — **iterate.** Local server at `:8888`, but your decorated functions
  execute on **remote GPU/CPU workers**. Hot-reloads on save and **streams the worker's
  logs live** to the terminal. No build/upload/deploy wait — use this the whole time you
  develop.
- `flash deploy` — **ship.** Builds an artifact and deploys a stable endpoint. Slow
  (build + upload + provision); only do this once the code works under `flash dev`.

`flash dev` ships **only the function body** to the worker, so a `NameError` for a
module-level name surfaces immediately here. `flash deploy` imports the whole module and
can mask that bug (see Gotcha #1). Develop against `flash dev` and you catch it first.

## Autonomous Dev Loop

`flash dev` is a long-running server. Three rules:
- **Run it in the background** — don't block on it.
- **Capture its output** to a log file.
- **Drive it over HTTP.**

The captured log is the remote worker's live stream (cold start, model load, `print`s,
tracebacks) — read it to debug.

```bash
flash dev > /tmp/flash-dev.log 2>&1 &                          # background; never run it blocking
for i in $(seq 1 60); do grep -q "flash dev  localhost:" /tmp/flash-dev.log && break; sleep 2; done  # bounded ~2min; if it never appears, check the log for errors
URL=$(grep -o "localhost:[0-9]*" /tmp/flash-dev.log | head -1)               # actual port (8888 bumps if taken)
curl -s "$URL/main/predict" -d '{"data": {...}}'               # dispatches to the remote worker
```

- **Read the real URL from the log** — flash auto-bumps the port if 8888 is in use, and
  prints `✓ flash dev  localhost:<port>` plus the route table.
- **Routes are namespaced by file**: `main.py`'s `/predict` is served at `/main/predict`.
- **Two route shapes, two body shapes** (mismatch → `422` naming the missing field in `loc`):
  - **Load-balanced** (`@api.post("/predict")`) → `POST /main/predict`, body is the arg
    at top level: a handler `def predict(data: dict)` wants `{"data": {...}}` (not the bare object).
  - **Queue-based** (bare `@Endpoint` decorator) → `POST /main/runsync` (the local dev
    server only generates `/runsync`; production also exposes `/run`),
    body is **double-wrapped** in `input`: a handler `def synthesize(data: dict)` wants
    `{"input": {"data": {...}}}`. The outer `input` is the queue envelope; the inner key is
    the handler's param name.
- Edit a handler and save — hot-reload re-syncs the body; just re-send the request, no
  redeploy. Add `--auto-provision` to skip the first-call cold start. `kill %1` when done.

## Endpoint: Three Modes

Full constructor params and the GPU/CPU enum tables are in [reference/api.md](reference/api.md).

### Mode 1: Your Code (Queue-Based Decorator)

One function = one endpoint with its own workers.

```python
from runpod_flash import Endpoint, GpuGroup

@Endpoint(name="my-worker", gpu=GpuGroup.AMPERE_80, workers=5, dependencies=["torch"])
async def compute(data):
    import torch  # MUST import inside function (cloudpickle)
    return {"sum": torch.tensor(data, device="cuda").sum().item()}

result = await compute([1, 2, 3])
```

### Mode 2: Your Code (Load-Balanced Routes)

Multiple HTTP routes share one pool of workers.

```python
from runpod_flash import Endpoint, GpuGroup

api = Endpoint(name="my-api", gpu=GpuGroup.ADA_24, workers=(1, 5), dependencies=["torch"])

@api.post("/predict")
async def predict(data: list[float]):
    import torch
    return {"result": torch.tensor(data, device="cuda").sum().item()}

@api.get("/health")
async def health():
    return {"status": "ok"}
```

### Mode 3: External Image (Client)

Deploy a pre-built Docker image and call it via HTTP.

```python
from runpod_flash import Endpoint, GpuGroup, PodTemplate

server = Endpoint(
    name="my-server",
    image="my-org/my-image:latest",
    gpu=GpuGroup.AMPERE_80,
    workers=1,
    env={"HF_TOKEN": "xxx"},
    template=PodTemplate(containerDiskInGb=100),
)

# LB-style
result = await server.post("/v1/completions", {"prompt": "hello"})
models = await server.get("/v1/models")

# QB-style
job = await server.run({"prompt": "hello"})        # optional: webhook="https://..." for completion callback
await job.wait()
print(job.output)
```

Connect to an existing endpoint by ID (no provisioning):

```python
ep = Endpoint(id="abc123")
job = await ep.runsync({"prompt": "hello"})  # runsync wraps this as {"input": {"prompt": "hello"}}
print(job.output)
```

## How Mode Is Determined

| Parameters | Mode |
|-----------|------|
| `name=` only | Decorator (your code) |
| `image=` set | Client (deploys image, then HTTP calls) |
| `id=` set | Client (connects to existing, no provisioning) |

The table above is *how* the mode is picked from params. *When* to reach for `image=`:

### When to use `image=` (custom container) vs your own code

Default to writing Python (decorator / routes) — it runs arbitrary code with
`dependencies=[...]`/`system_dependencies=[...]` and needs no Dockerfile. Even large
HuggingFace models stay in decorator mode (weights stream at runtime — see
[reference/patterns.md → Loading ML models](reference/patterns.md#loading-ml-models-warm-workers)).
Reach for `image=` **only** when you need:

- **a pre-built inference server** — vLLM, TensorRT-LLM (`image="vllm/vllm-openai:latest"`, or `runpod/worker-vllm`, `runpod/worker-comfy`)
- **system-level deps not pip-installable** — a specific CUDA/cuDNN, OS libraries
- **models baked into the image** — to skip the runtime download entirely
- **an existing Runpod Serverless worker** — you already have a working image

Trade-off: `image=` mode **can't run arbitrary Python** (the image owns all logic) and the
image must implement a Runpod Serverless handler. Full list + examples:
https://docs.runpod.io/flash/custom-docker-images

## Gotchas

1. **Only the function body ships to the worker** -- most common error. Put imports *and* any module-level constants/helpers the function uses *inside* the decorated body. `flash deploy` imports the whole module so module globals happen to work; `flash dev` ships just the body, so a module-level name raises `NameError`. A handler that works deployed can break under dev — fix it by moving everything inside.
2. **Forgetting await** -- all decorated functions and client methods need `await`.
3. **Missing dependencies** -- must list in `dependencies=[]`.
4. **gpu/cpu are exclusive** -- pick one per Endpoint.
5. **idle_timeout is seconds** -- default 60s, not minutes.
6. **10MB payload limit** -- pass URLs, not large objects. Return binary (audio/images/files) as base64 in the JSON (`{"audio_b64": ...}`) and decode client-side; for larger outputs write to a NetworkVolume or upload to storage and return a URL.
7. **Client vs decorator** -- `image=`/`id=` = client. Otherwise = decorator.
8. **Auto GPU switching requires workers >= 5** -- pass a list of GPU types (e.g. `gpu=[GpuGroup.ADA_24, GpuGroup.AMPERE_80]`) and set `workers=5` or higher. The platform only auto-switches GPU types based on supply when max workers is at least 5.
9. **`runsync` timeout is 60s** -- cold starts can exceed 60s. Use `ep.runsync(data, timeout=120)` for first requests or use `ep.run()` + `job.wait()` instead.
10. **Request body shape (raw/external HTTP callers only)** -- match the request shape to the endpoint type:
    - **LB routes** (`@api.post(...)`): send the handler arg at the top level — `{"data": {...}}`.
    - **QB endpoints** (bare `@Endpoint`, hit via `.../run` or `.../runsync`): the worker calls
      **`handler(**job_input)`**, so the request's `input` keys must match the handler's parameter
      names — `def transcribe(input_data: dict)` wants `{"input": {"input_data": {...}}}`, and
      `def read(input: dict)` wants `{"input": {"input": {...}}}`. A mismatch fails with
      `got an unexpected keyword argument …`. Use `**kwargs` if the handler ignores the payload.
    - **Never send an empty `input`.** A QB request with `{"input": {}}` is rejected by the
      worker SDK as `Job has missing field(s): id or input` — always include at least one key.
    - *Context:* the flash client (`ep.runsync(x)`, `api.post(...)`) hides the spreading, so this
      only bites raw HTTP/external callers (mismatch behavior verified 2026-07-10 via worker logs).
      See *Autonomous Dev Loop*.
11. **Load a model once per worker (not per call)** -- for real inference use a class `@Endpoint` whose `__init__` loads the model once per worker (see [reference/patterns.md → Loading ML models](reference/patterns.md#loading-ml-models-warm-workers)). In function-form, reconcile with #1 by caching in a module global *inside* the body so it works under both `flash dev` and `deploy`:
    ```python
    global _MODEL
    try: _MODEL
    except NameError: _MODEL = load_model()   # runs once per worker, reused across calls
    ```
12. **Native CUDA libs go in `dependencies=[]` too** -- e.g. CTranslate2/faster-whisper needs `nvidia-cublas-cu12` + `nvidia-cudnn-cu12` or it silently falls back to CPU. Add them alongside the Python package.
13. **Silent 401 auth failure** -- a set `RUNPOD_API_KEY` env var overrides the `flash login` token, so a bad/expired key wins. The failure is quiet: provisioning logs `GraphQL request failed: 401`, but `flash dev` still prints its normal ready line ("failed endpoints deploy on-demand"), so it *looks* healthy. When endpoints fail to provision:
    1. Check the provisioning log for `GraphQL request failed: 401`.
    2. Verify the current key independently: `curl -s -o /dev/null -w '%{http_code}' https://rest.runpod.io/v1/endpoints -H "Authorization: Bearer $RUNPOD_API_KEY"` (200 = good, 401 = bad).
    3. Fix it: `unset RUNPOD_API_KEY` to fall back to the `flash login` token, or `export` a valid key.
14. **`system_dependencies=` adds to cold start** -- apt packages (e.g. `["ffmpeg", "espeak-ng"]`) install on the worker before first use, so the initial call is slower (on top of any model download); warm calls are unaffected.
15. **Teardown a deployed app with `flash app delete <app>`** -- `flash undeploy list` may show "no endpoints" for an app that is deployed and serving; `flash app delete` (or `runpodctl serverless delete <id>`) reliably removes it.

## Resources

- Setup & CLI: [reference/setup-and-cli.md](reference/setup-and-cli.md) · API & compute enums: [reference/api.md](reference/api.md) · Patterns: [reference/patterns.md](reference/patterns.md)
- Flash source: https://github.com/runpod/flash
- Runnable examples: https://github.com/runpod/flash-examples — clone and adapt the closest one
- Package (PyPI): https://pypi.org/project/runpod-flash/
- Docs: https://docs.runpod.io/flash/overview
  - Custom Docker images (when + how): https://docs.runpod.io/flash/custom-docker-images
  - Storage / network volumes: https://docs.runpod.io/flash/configuration/storage

Referenced files: 10

runpod10.9 KB

View saved version →

---
name: runpod
description: >-
  Start here for any Runpod task — running GPU/CPU pods, deploying serverless
  endpoints, templates, network volumes, building images, or understanding how
  Runpod works. Routes the request to the right Runpod skill (runpod-mcp,
  runpodctl, flash, companion-clis, or runpod-usage). Use when it is unclear
  which Runpod skill applies.
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Runpod (router)

The entrypoint for the Runpod skills. This skill does no work itself — it picks
the right lane and hands off. Read the matching skill's `SKILL.md` next.

## The lanes

| Lane | Use it for |
| --- | --- |
| **runpod-mcp** | Manage infra (pods, endpoints, jobs, templates, volumes, registries, catalog, billing) via **structured tool calls** — when the Runpod MCP tools are connected in this session. |
| **runpodctl** | Manage the same infra from a **terminal/CI/script**, plus the things only the CLI does: Hub browse/deploy, `send`/`receive` file transfer, SSH keys, `doctor` setup, model cache. |
| **flash** | **Write Python** that runs on Runpod serverless — `@remote`/`@Endpoint` functions, `flash dev` hot-reload, `flash deploy`. Code-first, not infra management. |
| **companion-clis** | **Prerequisite artifacts**: download a model (`hf`), build/push an image (`docker`), repos/releases (`gh`), move data to a network volume over S3 (`aws`). |
| **runpod-usage** | **Understand** how Runpod works before acting — pods vs serverless, building a container, storage, GPU selection, gotchas. Knowledge only. |

## First run — check auth before the first infra action

Infra tasks (pods, endpoints, jobs, volumes) need a working control plane — the **Runpod MCP**
or **runpodctl**. Don't start and discover mid-task that nothing's set up: check first, and if
it isn't, help the user set up rather than limping on a partial fallback.

**Check** (credential resolution order: `RUNPOD_API_KEY` env → `.env` → `~/.runpod/config.toml`):
```bash
runpodctl user            # succeeds ⇒ a key is set and valid
```
Plus, in Claude Code, `/mcp` should show `runpod` **Connected**.

**Rule: get a key first — do not default to MCP OAuth.** The reason: one `RUNPOD_API_KEY`
unlocks every tool — it authenticates **runpodctl + flash + the hosted MCP** (as `--header
"Authorization: Bearer $RUNPOD_API_KEY"`). The MCP's "Sign in with Runpod" OAuth auths the **MCP alone** — the CLIs
stay blocked, so you hit a wall on any CLI-only task (Hub, `send`/`receive`, SSH, `doctor`,
model cache/Model Repository, CPU endpoints). ⚠️ **OAuth-only is a half-setup.** If nothing's
set up, stop and get a key, in order:
1. **`flash login`** — browser OAuth that saves a real key to `~/.runpod/config.toml` (runpodctl
   + flash read it; reuse it for the MCP Bearer). One step, unlocks all. Human-only.
2. **`export RUNPOD_API_KEY=…`** (https://console.runpod.io/user/settings) — same full unlock;
   best for headless agents.
3. **MCP OAuth only** (`/mcp` → *Sign in*) — last resort, MCP-only work; CLIs stay unauthed.

**Then:** if a lane already works, use it — but if *only* the MCP is OAuth'd, still get a key
before any CLI-only step. Missing a CLI? `curl -sSL https://cli.runpod.net | bash` (runpodctl) ·
`uv tool install runpod-flash` (flash) · `npx @runpod/mcp-server@latest add` (MCP). Full setup:
[`runpod-usage/reference/getting-started.md`](../runpod-usage/reference/getting-started.md).

## How to route

1. **Conceptual question, or an unmade design choice** (serverless vs pod? which
   GPU? bake the model or mount a volume?) → read **runpod-usage** first, then
   continue with the answer.
2. **Write/iterate/ship your own code on Runpod GPUs** → **flash**.
3. **Produce an artifact** (download a model, build+push an image, create a repo
   release, sync data to a volume) → **companion-clis**.
4. **Manage infrastructure** (create/list/update/delete pods, endpoints,
   templates, volumes; list GPUs/data centers; run a serverless job; billing):
   - Capability only the CLI has — **Hub, `send`/`receive`, SSH keys, `doctor`,
     model cache** → **runpodctl**.
   - Otherwise, if the Runpod **MCP tools are connected** in this session
     (`create-pod`, `list-endpoints`, … are available) → **runpod-mcp**.
   - Otherwise (shell-only agent, no MCP) → **runpodctl**.

### runpod-mcp vs runpodctl (the overlap)

Both drive the same Runpod API, so they overlap on infra CRUD. Choose by
**capability first, environment second**:

- **MCP wins on convenience** for simple, structured operations — reads and basic
  CRUD — when its tools are connected (typed params, no shell quoting).
- **runpodctl takes over when an operation needs a capability MCP lacks** — even
  if MCP is connected — and is the only option for a shell-only agent or when the
  user wants a reproducible command.

Capability matrix (pick the preferred lane per operation):

| Operation | Preferred lane | Why |
| --- | --- | --- |
| List/get anything; start/stop/restart/delete a pod; simple CRUD on endpoints, templates, volumes, registries; catalog; billing | **runpod-mcp** if connected, else runpodctl | Simple structured ops — MCP is typed and convenient |
| Create a **simple** pod (one image + one GPU) | **runpod-mcp** if connected, else runpodctl | Both handle it |
| Create a pod **from a template** or a **CPU** pod | **runpod-mcp** if connected, else runpodctl | MCP's create-pod takes `templateId` (v2-only) and `computeType: "CPU"` |
| Create a pod with a **multi-GPU priority list**, or **template + CPU together** | **runpodctl** | MCP narrows to one GPU type, and rejects a template deploy for a CPU pod |
| Deploy from the **Hub** | **runpod-mcp** if connected, else runpodctl | MCP has `list-hub-repos` + `deploy-hub-repo` |
| **File transfer** (`send`/`receive`), **SSH** keys/info, **`doctor`** setup, **model** cache | **runpodctl** | MCP has no tool for these |
| Invoke a serverless job (`run`/`runsync`/status/stream) | **runpod-mcp** if connected, else runpodctl | MCP has first-class job tools |

Rule of thumb: **default to MCP for the easy stuff, hand off to runpodctl the
moment an op needs a flag/feature MCP doesn't expose.**

## Deploying a workload (the golden loop)

For any "get <X> running on Runpod" task, follow the **development loop** in
`runpod-usage/reference/development-loop.md`: decide pod vs serverless → provision → set up
(only if from-scratch) → verify → deliver → cost-guard + teardown. Two rules bind within it:

- **Prefer a prebuilt template / Hub worker over building an image from scratch.**
- **Before delivering, verify the workload with a real request from outside the pod/endpoint
  — a "Running"/"ready" status does not mean it is serving.**

It branches to two sub-loops:

- **Service you open at a URL** (Ollama, ComfyUI, dev box) →
  [`runpod-usage/reference/pod-workflows.md`](../runpod-usage/reference/pod-workflows.md)
  (ports + env + volume at creation, SSH-exec install, bind `0.0.0.0`, poll the
  proxy URL). Execute in the runpodctl lane.
- **Request/response API that scales to zero** (Whisper, inference) →
  [`runpod-usage/reference/endpoint-workflows.md`](../runpod-usage/reference/endpoint-workflows.md)
  (Hub worker vs flash vs custom image; invoke `/run`/`/runsync`; poll job status).

## Worked examples (golden paths)

**Two dozen** end-to-end scenarios (nearly all **live-verified**) live in
[`./golden-paths/README.md`](./golden-paths/README.md) — the yardstick for "can
an agent finish the job", with real commands + observed output to copy from. When a
task matches one, **open its golden path first** instead of re-deriving it:

| Want to… | Golden path |
| --- | --- |
| Run a server (Ollama/ComfyUI) on a pod at a URL | [01](./golden-paths/01-ollama-pod.md), [02](./golden-paths/02-comfyui-pod/README.md) |
| Deploy a serverless model endpoint (Hub / flash / custom image) | [03](./golden-paths/03-whisper-endpoint/README.md), [05](./golden-paths/05-model-to-endpoint-pipeline.md) |
| Serve a HuggingFace model without baking it in or a volume (host-cached) | [20 — model caching (`--model-reference`)](./golden-paths/20-model-caching-endpoint.md) |
| Call a ready hosted model (no deploy) | [11 — Public Endpoints](./golden-paths/11-public-endpoints.md) |
| Fine-tune, then serve the result | [04](./golden-paths/04-finetune-pod.md), [08](./golden-paths/08-finetune-to-serverless.md) |
| Interactive dev box (SSH / VS Code) | [06](./golden-paths/06-dev-pod.md) |
| Move data pod → volume → serverless | [07](./golden-paths/07-network-volume-handoff.md) |
| **Custom serverless when flash isn't enough** (dual-mode image dev loop) | [09](./golden-paths/09-custom-serverless-dev-loop/README.md) |
| Build a minimal image for a target (pod vs serverless queue) | [22 (pod)](./golden-paths/22-minimal-pod-image/README.md), [23 (queue)](./golden-paths/23-minimal-queue-image/README.md); concepts in [building-images](../runpod-usage/reference/building-images.md) |
| Decide what to bake into the image vs mount on a network volume | [25 — bake vs mount](./golden-paths/25-bake-vs-mount/README.md) |
| **High availability / multi-region** serverless (multi-volume + data sync) | [10](./golden-paths/10-multi-region-ha-serverless.md), [19 (3-region)](./golden-paths/19-three-region-same-file.md) |
| Stream output incrementally (`/stream`) | [12](./golden-paths/12-serverless-streaming.md) |
| Tune autoscaling / raise per-worker throughput | [13 (autoscaling)](./golden-paths/13-autoscaling-tuning.md), [18 (concurrency)](./golden-paths/18-concurrent-handler.md) |
| Load-balancing / HTTP-server or WebSocket worker | [14 (LB)](./golden-paths/14-load-balancing-endpoint.md), [17 (WebSocket)](./golden-paths/17-serverless-websocket.md) |
| Get notified on job completion (push, not poll) | [16 — webhooks](./golden-paths/16-serverless-webhooks.md) |
| Check health / debug a failing endpoint | [15 — monitor & debug](./golden-paths/15-monitor-and-debug.md) |

## Multi-lane tasks

Sequence is always **understand → produce artifacts → manage infra → verify**,
because infra can only reference artifacts that already exist. Keep each step in
one lane, and switch lanes at credential boundaries.

Example — "deploy `openai/gpt-oss-20b` to a serverless endpoint":
1. **runpod-usage** — serverless vs pod, GPU tier for 20B, bake vs mount vs cache.
2. **companion-clis** — `hf download …`, `docker build --platform=linux/amd64 …`, `docker push`.
3. **runpod-mcp** or **runpodctl** — create the endpoint referencing the image + GPU pool.
4. Same infra lane — invoke the endpoint / check status to verify.

## Auth

Everything is one key: **`RUNPOD_API_KEY`** (https://console.runpod.io/user/settings).
Each lane just makes that key resolvable — `runpodctl doctor`, `flash login`, MCP
stdio env var, or MCP hosted "Sign in with Runpod" (OAuth, no key on disk).
Companion CLIs use their **own** credentials (HuggingFace token, GitHub auth,
Docker Hub PAT, Runpod **S3** keys for `aws`) — do not reuse `RUNPOD_API_KEY` for
those.

Referenced files: 43

runpodctl20.7 KB

View saved version →

---
name: runpodctl
description: >-
  Runpod CLI for managing GPU/CPU workloads from the terminal — pods, serverless
  endpoints, templates, network volumes, Hub deploys, models, SSH, and file
  transfer (send/receive). Use for terminal/CI/scripting, Hub browse/deploy, SSH
  setup, `doctor`, or when the Runpod MCP tools are not connected. For structured
  tool calls in an MCP-enabled session, prefer runpod-mcp.
allowed-tools: Bash(runpodctl:*)
compatibility: Linux, macOS
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Runpodctl

Manage GPU pods, serverless endpoints, templates, volumes, and models.

## Install

`curl -sSL https://cli.runpod.net | bash` (any platform) or `brew install runpod/runpodctl/runpodctl`. Manual binaries, Windows/Linux steps, and the version caveat (`--model-reference` + multi-volume need **v2.4.0+**): **[reference/install.md](reference/install.md)**.

> Old runpodctl builds silently lack newer flags/behaviors (e.g. `--model-reference` doesn't
> exist before v2.4.0) and produce confusing downstream errors — and the Homebrew tap can lag
> well behind. So, before any work:
>
> - **Update to the latest build** — check `runpodctl version`, then run `runpodctl update`
>   (or reinstall from the [latest release](https://github.com/runpod/runpodctl/releases)).
> - **Pin to one recent version for the whole task.**
> - **Never switch between an old and a new binary mid-task** (that flip-flop is a known failure).
> - **Verify once** — `runpodctl version` shows the current build before you continue.

## Quick start

```bash
runpodctl update                    # FIRST: get on the latest build — old versions cause confusing errors
runpodctl version                   # confirm the current version before doing any work
export RUNPOD_API_KEY=your_key      # Non-interactive auth (agents) — runpodctl reads this
runpodctl doctor                    # Interactive first-time setup (API key + SSH) — for humans
runpodctl --help                    # See current top-level commands
runpodctl pod create --help         # Inspect exact current flags before creating
runpodctl gpu list                  # See available GPU types
runpodctl datacenter list           # GPU availability per data center (use to co-locate GPU + volume)
runpodctl hub search vllm           # Find a hub repo
runpodctl serverless create --hub-id <id> --name "my-vllm"  # Deploy from hub
runpodctl template search pytorch   # Find a template
runpodctl pod create --template-id runpod-torch-v21 --gpu-id "NVIDIA GeForce RTX 4090"  # Create from template
runpodctl pod list                  # List your pods
```

> Auth: an agent should `export RUNPOD_API_KEY=...` (non-interactive). `runpodctl
> doctor` is interactive (prompts) and also sets up SSH keys — good for a human's
> first run, not for scripted use.

API key: https://console.runpod.io/user/settings

## Live Help Is Authoritative

Live `runpodctl --help` output is authoritative for exact flags, aliases, and command syntax. Use this skill for workflows, decision rules, safety notes, and common examples.

```bash
runpodctl --help
runpodctl <resource> --help
runpodctl <resource> <action> --help
```

Before using unfamiliar commands, inspect live help first. Do not rely on this skill as an exhaustive flag reference.

**What live help does *not* cover:** output shapes, error codes, and exit-code behavior. `--help` lists flags; it never shows you what a failure looks like. For those, use [reference/output-and-errors.md](reference/output-and-errors.md) — and when in doubt, **probe the binary**: run the command wrong on purpose (`runpodctl serverless get nope`) and read the JSON it emits. Every doc is a snapshot, this skill included; the binary in front of you wins.

## Output & errors

Data is **JSON on stdout** (`--output=yaml` is the only alternative — there is no table
format; anything else silently returns JSON). A failure from the resource commands is a
single flat JSON object on **stderr** plus a **non-zero exit**:

```jsonc
{"error":"failed to get endpoint: endpoint not found","code":"not_found","status":404}
```

**Branch on `code`, never on `status` or the message.** `status` is there only when the
failure arrived on a non-2xx response — GraphQL reports a missing resource as HTTP 200 +
null data, so `if status == 404` misses every GraphQL not-found.

| `code` | what to do |
| --- | --- |
| `network_error` | **retry with backoff** — the only code meaning "couldn't reach the API" |
| `rate_limited` `server_error` | **retry with backoff** — 429/5xx from the API |
| `usage_error` `cli_error` `bad_request` `not_found` `conflict` | don't retry, fix the input |
| `no_credentials` | no key set: `export RUNPOD_API_KEY=…` or `runpodctl doctor` |
| `unauthorized` `forbidden` | a key **is** set but is wrong/expired or lacks access — don't retry, don't re-prompt for a missing key |
| anything else | treat as fatal, surface `error` verbatim — the API can pass through its own code |

**runpodctl never retries internally**; nothing backs off for you.

- **`not_found` always means the API lacks the resource**, never a mistyped local path
  (that's `cli_error`).
- **`cli_error` is a mixed bucket:** local environment problems *and* invocation mistakes
  the command validates itself (e.g. `ssh remove-key` with neither `--name` nor
  `--fingerprint`). Only cobra-enforced required flags are `usage_error`.
- **`usage_error`** = unknown command/flag, bad args, missing cobra-required flag; usage
  text follows the JSON. Runtime errors no longer print usage.
- **Non-empty stderr does not mean failure** — deprecation `warning:` and `note:` lines go
  to stderr on success too. Gate on the exit code, then parse stderr.

Coded errors, the serverless `urls` object and GPU pricing all need **runpodctl ≥ v2.8.0**.
Older binaries emit `{"error":"…"}` with **no `code` and no `status`** — still JSON-shaped,
so a `switch (err.code)` silently gets `undefined` rather than failing loudly. **Gate on
`code` being present**, not on JSON-vs-plaintext; `runpodctl version` is unreliable for
this (plaintext, and a source build reports a placeholder version).

Full code table, the surfaces that still print plaintext (`exec`, legacy `pod`
commands, `project`), and the env-var table (incl. `RUNPOD_INVOKE_URL`):
**[reference/output-and-errors.md](reference/output-and-errors.md)**.

## Decision Rules

- Use Hub when the user wants a known deployable app or worker such as vLLM, ComfyUI, Whisper, or a Runpod-maintained repo.
  - **Picking a worker:** prefer a **first-party or well-adopted, recently-released** worker on a **broad, high-availability GPU pool**. Observable signals via `runpodctl hub list`: `--owner runpod-workers` (first-party), `--order-by releasedAt`/`updatedAt` (recency), `--order-by deploys`/`stars` (adoption). Don't pin a scarce large-GPU tier a small model doesn't need.
- **"Active worker" = minimum workers, not maximum.** If a user asks for an "active worker," they mean `--workers-min 1` (keep one worker always warm → no cold start), **not** `--workers-max 1` (that only caps the ceiling). A warm min-1 worker is ideal for development/iteration.
- ⚠️ **A min-1 worker bills continuously, even while idle** (it defeats scale-to-zero). When you set `--workers-min 1` for dev, you **must** set it back to `--workers-min 0` (or delete the endpoint) when done — otherwise it quietly runs up cost.
- `serverless update` has **no `--gpu-id` flag**. To change an existing endpoint's GPU pool, call `PATCH https://rest.runpod.io/v1/endpoints/<id>` with `{"gpuTypeIds":[...]}` directly.
- **CPU serverless endpoints:** always create them with `runpodctl serverless create --compute-type CPU` — **not** the MCP server, whose v2 `create-endpoint` requires `gpuPoolIds` and has no CPU concept. **Never** use the public control REST `POST https://rest.runpod.io/v1/endpoints` with `"computeType":"CPU"` — it silently provisions a **GPU** endpoint instead (verified evidence in the Serverless command section below).
- Use templates when the user already has a template ID, wants reusable image/config defaults, or needs lower-level control than Hub.
- Use direct pod creation with `--image` when the user has a specific Docker image and does not need a saved template.
- Use serverless for request/response inference APIs and scalable workers; use pods for interactive work, notebooks, training, debugging, or long-lived sessions.
- Use CPU pods for preprocessing, file movement, lightweight scripts, and non-CUDA work. Use GPU pods when CUDA, model inference, training, or GPU memory is required.
- Do not pass GPU flags when creating CPU pods. Check `runpodctl pod create --help` for the current valid flag set.
- Standing up a **service on a pod** (Ollama, ComfyUI, a dev server)? Declare its `--ports` and `--env` **at creation** (they can't be added to a running pod without a reset), then follow the pod development loop in the `runpod-usage` skill (`reference/pod-workflows.md`) — SSH-exec the install, bind to `0.0.0.0`, and poll the proxy URL until it answers.
- For SSH, use `runpodctl pod get <pod-id>` or `runpodctl ssh info <pod-id>` to retrieve connection details. runpodctl has **no interactive-shell command** — `ssh info` returns the connection command + key but does not connect. Run commands over SSH yourself with `ssh user@host "command"`.
- Network volumes are location-sensitive. Check datacenter availability before attaching volumes, and use `send` / `receive` or S3-compatible storage for migrations.
- Clean up paid resources after tests: delete serverless endpoints, pods, and temporary volumes created for validation.
  - **Cost guard on creation:** use `--terminate-after` (deletes the pod); `--stop-after` only *stops* it, so disk/volume keep billing.
  - **Attached volume:** to delete a network volume, remove the pod using it first.

### Serverless facts (context, not rules)

- **Scale-to-zero billing:** serverless endpoints scale to zero with `--workers-min 0` (the default) — no GPU billing while idle, only per request-second; this is the right cost posture for a request/response API.
- **Broken-image tell:** if deployed workers go `ready` but jobs sit `IN_QUEUE` with `inProgress: 0`, the image is broken/mis-dispatching — the fix is to switch to a different worker rather than wait it out.
- **Diagnosing it:** there's no first-class serverless worker-log command, so diagnosis relies on `/health` worker counts.

## Commands

Essentials below. **Full flag menu → [reference/command-reference.md](reference/command-reference.md)** (pods lifecycle, hub/template filters, registry auth, billing, SSH key management); live `runpodctl <resource> <action> --help` is authoritative for exact flags.

### Pods

```bash
runpodctl pod list                                   # running pods (+ --all / --status / --since / --created-after)
runpodctl pod get <pod-id>                           # details incl. SSH info
runpodctl pod create --template-id <id> --gpu-id "NVIDIA GeForce RTX 4090"   # from template
runpodctl pod create --image <img> --gpu-id "NVIDIA GeForce RTX 4090"        # from image
runpodctl pod create --compute-type cpu --image ubuntu:22.04                 # CPU pod (lowercase `cpu`; serverless uses `CPU`)
runpodctl pod {start|stop|restart|reset|update|delete} <pod-id>              # lifecycle (delete aliases: rm/remove)
```

### Hub

Browse/search the Runpod Hub (curated deployable repos).

```bash
runpodctl hub search vllm                            # find a repo (+ hub list [--type/--category/--order-by/--owner])
runpodctl hub get <listing-id|owner/name>            # repo details
```

### Serverless (alias: sls)

```bash
runpodctl serverless list | get <endpoint-id> | delete <endpoint-id>
runpodctl serverless create --name "x" --template-id <id>       # from template
runpodctl serverless create --name "x" --hub-id <listing-id>    # from hub (+ --env KEY=VAL to override defaults)
runpodctl serverless create --hub-id <id> --gpu-id "NVIDIA GeForce RTX 4090" \
  --model-reference https://huggingface.co/<org>/<model>:main   # attach & host-cache a HF model (GPU only)
runpodctl serverless update <endpoint-id> --workers-max 5
```

**Invoke URLs come back with the endpoint.** `create`/`get`/`list`/`update` include a
`urls` object (`run`, `runsync`, `health`), so a freshly created endpoint is callable
without a second lookup — read them instead of assembling the URL yourself. They're
built from `RUNPOD_INVOKE_URL` (default `https://api.runpod.ai/v2`), which
`RUNPOD_API_URL`/`RUNPOD_GRAPHQL_URL` do **not** move: [reference/output-and-errors.md](reference/output-and-errors.md#serverless-invoke-urls).

**Create from hub:** `--hub-id` resolves the hub listing, extracts the build image and config (GPU IDs, container disk, env vars), creates an inline template, and deploys. Accepts both SERVERLESS and POD listing types. GPU IDs and env var defaults from the hub config are included automatically; override with `--gpu-id` and `--env`.

**CPU serverless endpoints** (the always/never rule is in Decision Rules above): create with `runpodctl serverless create --compute-type CPU` (optionally `--instance-id`, e.g. `cpu3g-4-16`). Verified evidence for why the public REST must not be used: 2026-07-14, `POST https://rest.runpod.io/v1/endpoints` with `"computeType":"CPU"` silently returned a GPU endpoint (`gpuCount:1`, `cpuFlavorIds:null`), while `runpodctl --compute-type CPU` correctly returned `computeType:"CPU"` with `instanceIds:["cpu3g-4-16"]`. The MCP server is **not** an alternative here: its v2 `create-endpoint` requires `gpuPoolIds` and the v2 spec has no `computeType`/`cpuFlavor` field at all (verified 2026-07-29). The public control REST is v1-only (`rest.runpod.io/v2` just redirects to docs). The separate **runtime/invoke** API `https://api.runpod.ai/v2/<endpoint-id>/…` (health/run/runsync/openai) is a different v2 and works fine — the v1-vs-v2 caveat here is only about the **control/management** REST.

**Model cache (`--model-reference`):** Attach a Hugging Face model to the endpoint by full URL with a ref, e.g. `https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct:main`. Runpod caches it host-side in the standard HF cache dir (`/runpod-volume/huggingface-cache/hub/`), so the worker loads it directly — no bake, no volume. Repeatable; works with `--template-id`/`--hub-id`, GPU only, **runpodctl v2.4.0+**. Full mechanics + how it compares to baking / network volume / the Model Repository: **[reference/model-caching.md](reference/model-caching.md)**. Worked end-to-end: golden path [20 — model-caching endpoint](../runpod/golden-paths/20-model-caching-endpoint.md).

**Multi-region / high-availability (`--network-volume-ids`):** attach **multiple** network
volumes (one per data center) so workers spread across DCs instead of being pinned to one —
`runpodctl serverless create --template-id <t> --network-volume-ids <v1>,<v2> --data-center-ids <dc1>,<dc2> …`.
**Requires runpodctl ≥ v2.4.0** (older versions don't support multi-volume attach). Check
`runpodctl version`; the Homebrew tap can lag, so prefer the
[GitHub releases](https://github.com/runpod/runpodctl/releases) binary. Data does **not**
sync between volumes automatically — see golden path
[10 — multi-region HA serverless](../runpod/golden-paths/10-multi-region-ha-serverless.md).

For exact serverless flags, run `runpodctl serverless <action> --help`.

### Templates (alias: tpl)

```bash
runpodctl template search <q>                        # find (+ template list [--type official/community/user, --all, --limit])
runpodctl template get <template-id>                 # details (README, env, ports)
runpodctl template create --name "x" --image "img" [--serverless]
runpodctl template delete <template-id>
```

### Network Volumes (alias: nv)

```bash
runpodctl network-volume list                         # List all volumes
runpodctl network-volume get <volume-id>              # Get volume details
runpodctl network-volume create --name "x" --size 100 --data-center-id "US-GA-1"  # Create volume
runpodctl network-volume update <volume-id> --name "new"  # Update volume
runpodctl network-volume delete <volume-id>           # Delete volume
```

For exact network volume flags, run `runpodctl network-volume <action> --help`.

> **No storage-tier flag.** `create` provisions the data center's **default** tier — there's
> no `--type`. To get a **High-Performance** volume, use the console (a ⚡ data center's toggle)
> or a raw **v2 REST** call (`POST https://v2-rest.runpod.io/v2/network-volumes` with
> `"type":"HIGH_PERFORMANCE"`) — or the MCP `create-network-volume` tool, which takes
> `volumeType` (`STANDARD` | `HIGH_PERFORMANCE`). Tier is immutable after creation. Launch details: golden path [21](../runpod/golden-paths/21-storage-tiers.md).

### Models (Model Repository)

`runpodctl model` manages the **Runpod Model Repository** — managed, versioned storage
for your **own** model artifacts (upload once, distributed to workers; not pinned to a
data center like a network volume). What it is, why/how, migrating off a baked-in model,
and Model-Repo-vs-volume: **[reference/model-caching.md](reference/model-caching.md)**.

```bash
runpodctl model list                                  # List your models
runpodctl model list --all                            # List all models (not just yours)
runpodctl model list --name "llama"                   # Filter by name
runpodctl model list --provider "meta"                # Filter by provider
runpodctl model add --name "my-model" --model-path ./model   # Upload a local model dir (multipart)
runpodctl model remove --name "my-model" --owner <owner>     # Remove a model
```

`model add` supports upload sessions, versioning, metadata, and private-source credentials — see live `runpodctl model add --help`.

### Info & SSH

```bash
runpodctl user                                       # account info + balance (alias: me)
runpodctl gpu list                                   # available GPUs + $/hr + per-DC stock (+ --include-unavailable)
runpodctl datacenter list                            # datacenters (alias: dc)
runpodctl ssh info <pod-id>                          # SSH connection details (command + key; NOT an interactive session)
```

**`gpu list` carries pricing and placement data** — `securePricePerHr` /
`communityPricePerHr` (explicitly `null` when that cloud doesn't offer the GPU) and a
`dataCenterAvailability[]` breakdown. Read the breakdown, not just top-level
`stockStatus` (which is only the *best* status across DCs), when a create has to
schedule in a specific DC — and pass `--include-unavailable`, since the default listing
hides no-stock GPUs and can omit one that has stock only in the DC you want. The prices
are **pod on-demand** rates. Shape, stock-value vocabulary and the `"none"` vs
omitted-key sentinel:
[reference/output-and-errors.md](reference/output-and-errors.md#gpu-pricing-and-per-data-center-availability).

`ssh info` gives connection details, not a session — if interactive SSH isn't available, run `ssh user@host "command"`. **Registry auth, `billing` history, and SSH key management** (`ssh add-key`/`remove-key`) are in [reference/command-reference.md](reference/command-reference.md).

### File Transfer

```bash
runpodctl send <path>                                # prints a one-time code, then blocks until the receiver connects
runpodctl receive <code>                             # positional code (no --code flag)
```

Encrypted/incremental/compressed — don't pre-tar. **Key gotchas:** capture the **first line of `send` stdout** (the code) as it streams (background + tee), each `send` mints a **fresh** code, both sides must exit `0`. Full agent flow (pod push via `ssh` + `receive`): [reference/command-reference.md](reference/command-reference.md#file-transfer).

### Utilities

```bash
runpodctl doctor                                      # Diagnose and fix CLI issues
runpodctl update                                      # Update CLI
runpodctl version                                     # Show version
runpodctl completion                                  # Auto-detect shell and install completion
```

## URLs

### Pod URLs

Access exposed ports on your pod:

```
https://<pod-id>-<port>.proxy.runpod.net
```

Example: `https://abc123xyz-8888.proxy.runpod.net`

### Serverless URLs

```
https://api.runpod.ai/v2/<endpoint-id>/run        # Async request
https://api.runpod.ai/v2/<endpoint-id>/runsync    # Sync request
https://api.runpod.ai/v2/<endpoint-id>/health     # Health check
https://api.runpod.ai/v2/<endpoint-id>/status/<job-id>  # Job status
```

`serverless create`/`get`/`list`/`update` already return `run`/`runsync`/`health` in a
`urls` object — prefer those over hand-assembling, since a non-default
`RUNPOD_INVOKE_URL` changes the base. Only `status/<job-id>` has to be built by hand.

## Source & docs

- CLI source: https://github.com/runpod/runpodctl
- Releases (binaries): https://github.com/runpod/runpodctl/releases
- Docs: https://docs.runpod.io/runpodctl/overview

Referenced files: 13

runpod-mcp7.56 KB

View saved version →

---
name: runpod-mcp
description: >-
  Manage Runpod infrastructure — pods, serverless endpoints, jobs, templates,
  network volumes, container-registry auth, GPU/CPU catalog, and billing — via
  the Runpod MCP server's structured tool calls. Use when the Runpod MCP tools
  (create-pod, list-endpoints, …) are connected in this session, or to connect
  them (hosted OAuth or local npx). Prefer this over runpodctl for plain infra
  CRUD when MCP is available; use runpodctl for the terminal, file transfer, or
  SSH setup.
allowed-tools: Bash(npx:*), Bash(claude:*)
compatibility: Linux, macOS, Windows
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Runpod MCP

The Runpod MCP server exposes Runpod's control plane as structured tool calls,
so an MCP-capable agent can manage infrastructure without shelling out. It is the
same Runpod REST API that `runpodctl` uses — pick MCP when its tools are
connected (typed params, structured errors, no shell quoting).

## Connect

Connect the hosted server with **your API key as a Bearer header** if you also use runpodctl/flash — that one key auths the MCP *and* the CLIs (the 80% path):

```bash
claude mcp add --transport http runpod -s user https://mcp.getrunpod.io/ \
  --header "Authorization: Bearer $RUNPOD_API_KEY"
```

Plain **OAuth** ("Sign in with Runpod", via `npx @runpod/mcp-server@latest add`) is MCP-only — the CLIs stay unauthed, so use it only for MCP-only work. Local **stdio** runs the server as a subprocess with your key. Those variants + the key-vs-OAuth tradeoff: **[reference/connect.md](reference/connect.md)**. After connecting, reconnect the client (in Claude Code, `/mcp`) so the tools load.

**Verify it's live (do this before relying on MCP):** in Claude Code run `/mcp` —
`runpod` should show **Connected**, not *Needs authentication* (if it's the latter,
sign in there first; the bundled plugin server registers the URL but stays inert
until you authenticate). Confirm a real call works by asking for `list-endpoints`.
If the `runpod` tools aren't present at all, the server isn't connected — (re)run the
install above, or fall back to **runpodctl** for this task.

**Check the server version (which REST API it drives):** the MCP `initialize` handshake
returns it in `serverInfo.version`. `/mcp` in Claude Code shows it, or probe the hosted
server directly:

```bash
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}' \
| curl -s -X POST https://mcp.getrunpod.io/ -H "Content-Type: application/json" \
    -H "Accept: application/json, text/event-stream" -H "Authorization: Bearer $RUNPOD_API_KEY" -d @-
# → serverInfo.version e.g. "3.0.0 [RUNPOD_REST_VERSION=v2]"  (verified 2026-07-29)
```

The MCP server drives Runpod's **REST v2** internally (`RUNPOD_REST_VERSION=v2`), so most
tools avoid the buggy **public `rest.runpod.io/v1`** control API. Two exceptions worth
knowing: the Hub, public-endpoint and `set-endpoint-gpus` tools go through GraphQL (so they
work under either REST version), and **CPU serverless endpoints are not creatable through
MCP** — v2 has no CPU-endpoint concept at all (`create-endpoint` requires `gpuPoolIds`), so
use `runpodctl serverless create --compute-type CPU` for those.

**Prefer MCP or `runpodctl` over hand-rolled `rest.runpod.io/v1` calls for creating endpoints.**

## Tool surface

Structured tools, grouped by resource:

- **Pods** — list, get, create, update, start, stop, restart, delete, stream logs.
- **Serverless endpoints** — list, get, create, update, delete; list workers; list releases; stream worker logs.
  - `create-endpoint` takes `endpointType: QUEUE` (default) or `LOAD_BALANCER` — see golden path 14. The routing type is fixed at creation; `update-endpoint` cannot change it.
  - Read an endpoint's invoke URLs from `requestUrls` on the get/list reply instead of assembling them.
  - To pin a specific GPU **SKU** on an existing endpoint use `set-endpoint-gpus`; `create-endpoint`/`update-endpoint` expose only `gpuPoolIds` and can't express a SKU (`deploy-hub-repo` can pin one at deploy time via `gpuIds` exclusions).
- **Jobs (serverless runtime)** — run, runsync, status, stream, cancel, retry, health, purge queue.
- **Hub** — `list-hub-repos` (public catalog of prebuilt Serverless workers and Pod templates: vLLM, ComfyUI, …) and `deploy-hub-repo`, which deploys a repo's listed release as an endpoint — the same as clicking Deploy on the Hub.
- **Public endpoints** — `list-public-endpoints`: managed pay-per-use model APIs (text/image/video/audio) that need no deployment. Call the returned endpointId with `run-endpoint`/`runsync-endpoint`.
- **Templates** — list, get, create, update, delete.
- **Network volumes** — list, get, create, update, delete. `create-network-volume` takes `volumeType` (`STANDARD` | `HIGH_PERFORMANCE`) and a size of 10–4096 GB; omit `volumeType` to get the data center's default tier. The tier is **immutable after creation** — `update-network-volume` can't change it.
- **Container registry auth** — list, get, create, delete. A username + password for **any** registry; pass the resulting id as `containerRegistryAuthId` on create-pod/create-endpoint.
- **ECR delegations** (`list-`/`create-`/`delete-registry-delegation`) — **AWS ECR only**, v2 only, and stores no credentials: you register a repository ARN and Runpod gets scoped pull access instead. Prefer it over a stored username/password for ECR. The reply carries a `dockerRegistryUri` — that's the image URI to deploy with.
- **Catalog** — list/get GPU types, list/get CPU types, list/get data centers.
- **Billing** — scoped usage/cost breakdowns (`get-billing`).

> The tool list above is a map, not a contract. The server is the source of truth —
> `/mcp` (or your client's tool list) shows exactly what the connected version exposes,
> and each tool carries its own parameter descriptions. Check there before assuming a
> capability exists or doesn't.

> Delete tools (`delete-template`, `delete-pod`, …) can return `isError: true` with
> "Unexpected end of JSON input" **even on success** — the Runpod REST API returns
> 204 No Content. Don't treat it as failure; confirm with a follow-up `get-`/`list-`
> (a deleted resource then 404s).

## Use MCP vs runpodctl

- **Use runpod-mcp** when the tools are connected AND the task is infra CRUD or a
  serverless job call the server exposes. Cap large job/log output to a file.
- **Use runpodctl instead** for: **`send`/`receive`** file transfer, **SSH** key
  management, **`doctor`** setup, **model cache** — or any shell-only agent, or
  when the user wants a reproducible command.
- **Hand pod creation to runpodctl** for a **multi-GPU priority list** (MCP's v2
  create-pod takes one GPU type; extra `gpuTypeIds` are dropped with a `_warning`
  on success), or for a **template + CPU** pod together — `create-pod` rejects that
  combination, since a template deploy is GPU-and-v2-only. Each alone is fine in
  MCP: `templateId` (v2-only, `imageName` then optional, and each field you pass
  replaces the template's whole value rather than merging) or `computeType: "CPU"`.
- **Not this lane:** writing/deploying your own Python (→ flash); downloading
  models or building/pushing images (→ companion-clis).

For concepts (pods vs serverless, GPU selection, storage), read
`../runpod-usage/`.

## Source & docs

- Server source: https://github.com/runpod/runpod-mcp
- Package (npm): https://www.npmjs.com/package/@runpod/mcp-server
- Hosted endpoint: https://mcp.getrunpod.io/
- Docs: https://docs.runpod.io

Referenced files: 2

runpod-usage2.26 KB

View saved version →

---
name: runpod-usage
description: >-
  How Runpod works and how to work it — pods vs serverless, GPU/VRAM selection,
  storage, building a container, networking, plus the agentic pod development loop
  (provision → ssh-exec → set up → poll readiness) and on-pod install hygiene
  (uv/apt). Use to answer "how does X work", "which GPU", "how do I build a
  container", or "how do I stand up a workload on a pod". Guidance, not a tool —
  execute with runpodctl, runpod-mcp, or flash.
metadata:
  author: runpod
  version: "1.1.2" # x-release-please-version
license: Apache-2.0
---

# Runpod usage (concepts)

Background knowledge for making the right choice before you act. This skill runs
nothing — once you know what to do, execute with **runpod-mcp**/**runpodctl**
(infra), **flash** (your own code), or **companion-clis** (models/images/data).

Read the one reference file that matches the question:

| Question | Read |
| --- | --- |
| First-run setup / auth — get + set `RUNPOD_API_KEY`, SSH, companion creds | `reference/getting-started.md` |
| Pods vs serverless, workers, cold starts, FlashBoot, queue vs load-balanced | `reference/concepts.md` |
| **The development loop for ANY workload (start here)** — plan → prefer prebuilt → provision → verify → teardown | `reference/development-loop.md` |
| **Stand up / iterate a workload on a pod** — the pod sub-loop | `reference/pod-workflows.md` |
| **Deploy / iterate a serverless endpoint** — Hub vs flash vs custom, invoke + verify | `reference/endpoint-workflows.md` |
| **Install software on a pod** — package hygiene, `uv`, non-interactive, caching | `reference/on-pod-setup.md` |
| Build a Docker image Runpod can run (handler contract, Dockerfile, `--platform=linux/amd64`) | `reference/docker.md` |
| **How to build an image well** — base image, layering, bake-in vs volume, pod vs serverless (queue/LB) contract | `reference/building-images.md` |
| Where data lives — container disk vs network volume, model caching, S3 access | `reference/storage.md` |
| Which GPU / how much VRAM / cost & availability / data centers | `reference/gpu-selection.md` |
| Reaching a pod or endpoint over HTTP (proxy URLs, exposed ports) | `reference/networking.md` |
| Common mistakes and how to avoid them | `reference/gotchas.md` |

Referenced files: 14

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
Apache-2.0
Package author
Runpod
Keywords
runpod, gpu, serverless, pods, mcp, runpodctl, flash, infrastructure

Declared capabilities

  • Skills
  • MCP
  • CLI

Some manifest fields differ or could not be read. The structured report retains the source references.

Package observed Sep 30, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 18:00 UTC
Collection status
Collected

Plugin_cd395a620ba481918f5e3b37ce9a123d

Download plugin data (JSON)