{"id":12775,"plugin_id":"Plugin_cd395a620ba481918f5e3b37ce9a123d","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:02:09.593Z","digest":"1202ae98021495d8a687c521e8ef75bea740f2313ab8ddf2000ce4fb48a6c730","against":null,"payload":{"name":"runpodctl","description":"Runpod CLI for managing GPU/CPU workloads from the terminal — pods, serverless endpoints, templates, network volumes, Hub deploys, models, SSH, and file transfer (send/receive). Use for terminal/CI/scripting, Hub browse/deploy, SSH setup, `doctor`, or when the Runpod MCP tools are not connected. For structured tool calls in an MCP-enabled session, prefer runpod-mcp.","included_files":[{"relative_path":"evals/cpu-pod-create.eval.md","size_in_bytes":503},{"relative_path":"evals/error-handling.eval.md","size_in_bytes":2165},{"relative_path":"evals/hub-deploy-serverless.eval.md","size_in_bytes":1533},{"relative_path":"evals/image-to-template-to-serverless.eval.md","size_in_bytes":1327},{"relative_path":"evals/invoke-urls-and-gpu-pricing.eval.md","size_in_bytes":2093},{"relative_path":"evals/pod-auto-terminate.eval.md","size_in_bytes":792},{"relative_path":"evals/pod-from-template-with-volume.eval.md","size_in_bytes":583},{"relative_path":"evals/pod-ssh-connect.eval.md","size_in_bytes":734},{"relative_path":"evals/serverless-autoscale-by-requests.eval.md","size_in_bytes":816},{"relative_path":"reference/command-reference.md","size_in_bytes":9116},{"relative_path":"reference/install.md","size_in_bytes":1459},{"relative_path":"reference/model-caching.md","size_in_bytes":5276},{"relative_path":"reference/output-and-errors.md","size_in_bytes":14509}],"skill_md_contents":"---\nname: runpodctl\ndescription: >-\n  Runpod CLI for managing GPU/CPU workloads from the terminal — pods, serverless\n  endpoints, templates, network volumes, Hub deploys, models, SSH, and file\n  transfer (send/receive). Use for terminal/CI/scripting, Hub browse/deploy, SSH\n  setup, `doctor`, or when the Runpod MCP tools are not connected. For structured\n  tool calls in an MCP-enabled session, prefer runpod-mcp.\nallowed-tools: Bash(runpodctl:*)\ncompatibility: Linux, macOS\nmetadata:\n  author: runpod\n  version: \"1.1.2\" # x-release-please-version\nlicense: Apache-2.0\n---\n\n# Runpodctl\n\nManage GPU pods, serverless endpoints, templates, volumes, and models.\n\n## Install\n\n`curl -sSL https://cli.runpod.net | bash` (any platform) or `brew install runpod/runpodctl/runpodctl`. Manual binaries, Windows/Linux steps, and the version caveat (`--model-reference` + multi-volume need **v2.4.0+**): **[reference/install.md](reference/install.md)**.\n\n> Old runpodctl builds silently lack newer flags/behaviors (e.g. `--model-reference` doesn't\n> exist before v2.4.0) and produce confusing downstream errors — and the Homebrew tap can lag\n> well behind. So, before any work:\n>\n> - **Update to the latest build** — check `runpodctl version`, then run `runpodctl update`\n>   (or reinstall from the [latest release](https://github.com/runpod/runpodctl/releases)).\n> - **Pin to one recent version for the whole task.**\n> - **Never switch between an old and a new binary mid-task** (that flip-flop is a known failure).\n> - **Verify once** — `runpodctl version` shows the current build before you continue.\n\n## Quick start\n\n```bash\nrunpodctl update                    # FIRST: get on the latest build — old versions cause confusing errors\nrunpodctl version                   # confirm the current version before doing any work\nexport RUNPOD_API_KEY=your_key      # Non-interactive auth (agents) — runpodctl reads this\nrunpodctl doctor                    # Interactive first-time setup (API key + SSH) — for humans\nrunpodctl --help                    # See current top-level commands\nrunpodctl pod create --help         # Inspect exact current flags before creating\nrunpodctl gpu list                  # See available GPU types\nrunpodctl datacenter list           # GPU availability per data center (use to co-locate GPU + volume)\nrunpodctl hub search vllm           # Find a hub repo\nrunpodctl serverless create --hub-id <id> --name \"my-vllm\"  # Deploy from hub\nrunpodctl template search pytorch   # Find a template\nrunpodctl pod create --template-id runpod-torch-v21 --gpu-id \"NVIDIA GeForce RTX 4090\"  # Create from template\nrunpodctl pod list                  # List your pods\n```\n\n> Auth: an agent should `export RUNPOD_API_KEY=...` (non-interactive). `runpodctl\n> doctor` is interactive (prompts) and also sets up SSH keys — good for a human's\n> first run, not for scripted use.\n\nAPI key: https://console.runpod.io/user/settings\n\n## Live Help Is Authoritative\n\nLive `runpodctl --help` output is authoritative for exact flags, aliases, and command syntax. Use this skill for workflows, decision rules, safety notes, and common examples.\n\n```bash\nrunpodctl --help\nrunpodctl <resource> --help\nrunpodctl <resource> <action> --help\n```\n\nBefore using unfamiliar commands, inspect live help first. Do not rely on this skill as an exhaustive flag reference.\n\n**What live help does *not* cover:** output shapes, error codes, and exit-code behavior. `--help` lists flags; it never shows you what a failure looks like. For those, use [reference/output-and-errors.md](reference/output-and-errors.md) — and when in doubt, **probe the binary**: run the command wrong on purpose (`runpodctl serverless get nope`) and read the JSON it emits. Every doc is a snapshot, this skill included; the binary in front of you wins.\n\n## Output & errors\n\nData is **JSON on stdout** (`--output=yaml` is the only alternative — there is no table\nformat; anything else silently returns JSON). A failure from the resource commands is a\nsingle flat JSON object on **stderr** plus a **non-zero exit**:\n\n```jsonc\n{\"error\":\"failed to get endpoint: endpoint not found\",\"code\":\"not_found\",\"status\":404}\n```\n\n**Branch on `code`, never on `status` or the message.** `status` is there only when the\nfailure arrived on a non-2xx response — GraphQL reports a missing resource as HTTP 200 +\nnull data, so `if status == 404` misses every GraphQL not-found.\n\n| `code` | what to do |\n| --- | --- |\n| `network_error` | **retry with backoff** — the only code meaning \"couldn't reach the API\" |\n| `rate_limited` `server_error` | **retry with backoff** — 429/5xx from the API |\n| `usage_error` `cli_error` `bad_request` `not_found` `conflict` | don't retry, fix the input |\n| `no_credentials` | no key set: `export RUNPOD_API_KEY=…` or `runpodctl doctor` |\n| `unauthorized` `forbidden` | a key **is** set but is wrong/expired or lacks access — don't retry, don't re-prompt for a missing key |\n| anything else | treat as fatal, surface `error` verbatim — the API can pass through its own code |\n\n**runpodctl never retries internally**; nothing backs off for you.\n\n- **`not_found` always means the API lacks the resource**, never a mistyped local path\n  (that's `cli_error`).\n- **`cli_error` is a mixed bucket:** local environment problems *and* invocation mistakes\n  the command validates itself (e.g. `ssh remove-key` with neither `--name` nor\n  `--fingerprint`). Only cobra-enforced required flags are `usage_error`.\n- **`usage_error`** = unknown command/flag, bad args, missing cobra-required flag; usage\n  text follows the JSON. Runtime errors no longer print usage.\n- **Non-empty stderr does not mean failure** — deprecation `warning:` and `note:` lines go\n  to stderr on success too. Gate on the exit code, then parse stderr.\n\nCoded errors, the serverless `urls` object and GPU pricing all need **runpodctl ≥ v2.8.0**.\nOlder binaries emit `{\"error\":\"…\"}` with **no `code` and no `status`** — still JSON-shaped,\nso a `switch (err.code)` silently gets `undefined` rather than failing loudly. **Gate on\n`code` being present**, not on JSON-vs-plaintext; `runpodctl version` is unreliable for\nthis (plaintext, and a source build reports a placeholder version).\n\nFull code table, the surfaces that still print plaintext (`exec`, legacy `pod`\ncommands, `project`), and the env-var table (incl. `RUNPOD_INVOKE_URL`):\n**[reference/output-and-errors.md](reference/output-and-errors.md)**.\n\n## Decision Rules\n\n- Use Hub when the user wants a known deployable app or worker such as vLLM, ComfyUI, Whisper, or a Runpod-maintained repo.\n  - **Picking a worker:** prefer a **first-party or well-adopted, recently-released** worker on a **broad, high-availability GPU pool**. Observable signals via `runpodctl hub list`: `--owner runpod-workers` (first-party), `--order-by releasedAt`/`updatedAt` (recency), `--order-by deploys`/`stars` (adoption). Don't pin a scarce large-GPU tier a small model doesn't need.\n- **\"Active worker\" = minimum workers, not maximum.** If a user asks for an \"active worker,\" they mean `--workers-min 1` (keep one worker always warm → no cold start), **not** `--workers-max 1` (that only caps the ceiling). A warm min-1 worker is ideal for development/iteration.\n- ⚠️ **A min-1 worker bills continuously, even while idle** (it defeats scale-to-zero). When you set `--workers-min 1` for dev, you **must** set it back to `--workers-min 0` (or delete the endpoint) when done — otherwise it quietly runs up cost.\n- `serverless update` has **no `--gpu-id` flag**. To change an existing endpoint's GPU pool, call `PATCH https://rest.runpod.io/v1/endpoints/<id>` with `{\"gpuTypeIds\":[...]}` directly.\n- **CPU serverless endpoints:** always create them with `runpodctl serverless create --compute-type CPU` — **not** the MCP server, whose v2 `create-endpoint` requires `gpuPoolIds` and has no CPU concept. **Never** use the public control REST `POST https://rest.runpod.io/v1/endpoints` with `\"computeType\":\"CPU\"` — it silently provisions a **GPU** endpoint instead (verified evidence in the Serverless command section below).\n- Use templates when the user already has a template ID, wants reusable image/config defaults, or needs lower-level control than Hub.\n- Use direct pod creation with `--image` when the user has a specific Docker image and does not need a saved template.\n- Use serverless for request/response inference APIs and scalable workers; use pods for interactive work, notebooks, training, debugging, or long-lived sessions.\n- Use CPU pods for preprocessing, file movement, lightweight scripts, and non-CUDA work. Use GPU pods when CUDA, model inference, training, or GPU memory is required.\n- Do not pass GPU flags when creating CPU pods. Check `runpodctl pod create --help` for the current valid flag set.\n- Standing up a **service on a pod** (Ollama, ComfyUI, a dev server)? Declare its `--ports` and `--env` **at creation** (they can't be added to a running pod without a reset), then follow the pod development loop in the `runpod-usage` skill (`reference/pod-workflows.md`) — SSH-exec the install, bind to `0.0.0.0`, and poll the proxy URL until it answers.\n- For SSH, use `runpodctl pod get <pod-id>` or `runpodctl ssh info <pod-id>` to retrieve connection details. runpodctl has **no interactive-shell command** — `ssh info` returns the connection command + key but does not connect. Run commands over SSH yourself with `ssh user@host \"command\"`.\n- Network volumes are location-sensitive. Check datacenter availability before attaching volumes, and use `send` / `receive` or S3-compatible storage for migrations.\n- Clean up paid resources after tests: delete serverless endpoints, pods, and temporary volumes created for validation.\n  - **Cost guard on creation:** use `--terminate-after` (deletes the pod); `--stop-after` only *stops* it, so disk/volume keep billing.\n  - **Attached volume:** to delete a network volume, remove the pod using it first.\n\n### Serverless facts (context, not rules)\n\n- **Scale-to-zero billing:** serverless endpoints scale to zero with `--workers-min 0` (the default) — no GPU billing while idle, only per request-second; this is the right cost posture for a request/response API.\n- **Broken-image tell:** if deployed workers go `ready` but jobs sit `IN_QUEUE` with `inProgress: 0`, the image is broken/mis-dispatching — the fix is to switch to a different worker rather than wait it out.\n- **Diagnosing it:** there's no first-class serverless worker-log command, so diagnosis relies on `/health` worker counts.\n\n## Commands\n\nEssentials below. **Full flag menu → [reference/command-reference.md](reference/command-reference.md)** (pods lifecycle, hub/template filters, registry auth, billing, SSH key management); live `runpodctl <resource> <action> --help` is authoritative for exact flags.\n\n### Pods\n\n```bash\nrunpodctl pod list                                   # running pods (+ --all / --status / --since / --created-after)\nrunpodctl pod get <pod-id>                           # details incl. SSH info\nrunpodctl pod create --template-id <id> --gpu-id \"NVIDIA GeForce RTX 4090\"   # from template\nrunpodctl pod create --image <img> --gpu-id \"NVIDIA GeForce RTX 4090\"        # from image\nrunpodctl pod create --compute-type cpu --image ubuntu:22.04                 # CPU pod (lowercase `cpu`; serverless uses `CPU`)\nrunpodctl pod {start|stop|restart|reset|update|delete} <pod-id>              # lifecycle (delete aliases: rm/remove)\n```\n\n### Hub\n\nBrowse/search the Runpod Hub (curated deployable repos).\n\n```bash\nrunpodctl hub search vllm                            # find a repo (+ hub list [--type/--category/--order-by/--owner])\nrunpodctl hub get <listing-id|owner/name>            # repo details\n```\n\n### Serverless (alias: sls)\n\n```bash\nrunpodctl serverless list | get <endpoint-id> | delete <endpoint-id>\nrunpodctl serverless create --name \"x\" --template-id <id>       # from template\nrunpodctl serverless create --name \"x\" --hub-id <listing-id>    # from hub (+ --env KEY=VAL to override defaults)\nrunpodctl serverless create --hub-id <id> --gpu-id \"NVIDIA GeForce RTX 4090\" \\\n  --model-reference https://huggingface.co/<org>/<model>:main   # attach & host-cache a HF model (GPU only)\nrunpodctl serverless update <endpoint-id> --workers-max 5\n```\n\n**Invoke URLs come back with the endpoint.** `create`/`get`/`list`/`update` include a\n`urls` object (`run`, `runsync`, `health`), so a freshly created endpoint is callable\nwithout a second lookup — read them instead of assembling the URL yourself. They're\nbuilt from `RUNPOD_INVOKE_URL` (default `https://api.runpod.ai/v2`), which\n`RUNPOD_API_URL`/`RUNPOD_GRAPHQL_URL` do **not** move: [reference/output-and-errors.md](reference/output-and-errors.md#serverless-invoke-urls).\n\n**Create from hub:** `--hub-id` resolves the hub listing, extracts the build image and config (GPU IDs, container disk, env vars), creates an inline template, and deploys. Accepts both SERVERLESS and POD listing types. GPU IDs and env var defaults from the hub config are included automatically; override with `--gpu-id` and `--env`.\n\n**CPU serverless endpoints** (the always/never rule is in Decision Rules above): create with `runpodctl serverless create --compute-type CPU` (optionally `--instance-id`, e.g. `cpu3g-4-16`). Verified evidence for why the public REST must not be used: 2026-07-14, `POST https://rest.runpod.io/v1/endpoints` with `\"computeType\":\"CPU\"` silently returned a GPU endpoint (`gpuCount:1`, `cpuFlavorIds:null`), while `runpodctl --compute-type CPU` correctly returned `computeType:\"CPU\"` with `instanceIds:[\"cpu3g-4-16\"]`. The MCP server is **not** an alternative here: its v2 `create-endpoint` requires `gpuPoolIds` and the v2 spec has no `computeType`/`cpuFlavor` field at all (verified 2026-07-29). The public control REST is v1-only (`rest.runpod.io/v2` just redirects to docs). The separate **runtime/invoke** API `https://api.runpod.ai/v2/<endpoint-id>/…` (health/run/runsync/openai) is a different v2 and works fine — the v1-vs-v2 caveat here is only about the **control/management** REST.\n\n**Model cache (`--model-reference`):** Attach a Hugging Face model to the endpoint by full URL with a ref, e.g. `https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct:main`. Runpod caches it host-side in the standard HF cache dir (`/runpod-volume/huggingface-cache/hub/`), so the worker loads it directly — no bake, no volume. Repeatable; works with `--template-id`/`--hub-id`, GPU only, **runpodctl v2.4.0+**. Full mechanics + how it compares to baking / network volume / the Model Repository: **[reference/model-caching.md](reference/model-caching.md)**. Worked end-to-end: golden path [20 — model-caching endpoint](../runpod/golden-paths/20-model-caching-endpoint.md).\n\n**Multi-region / high-availability (`--network-volume-ids`):** attach **multiple** network\nvolumes (one per data center) so workers spread across DCs instead of being pinned to one —\n`runpodctl serverless create --template-id <t> --network-volume-ids <v1>,<v2> --data-center-ids <dc1>,<dc2> …`.\n**Requires runpodctl ≥ v2.4.0** (older versions don't support multi-volume attach). Check\n`runpodctl version`; the Homebrew tap can lag, so prefer the\n[GitHub releases](https://github.com/runpod/runpodctl/releases) binary. Data does **not**\nsync between volumes automatically — see golden path\n[10 — multi-region HA serverless](../runpod/golden-paths/10-multi-region-ha-serverless.md).\n\nFor exact serverless flags, run `runpodctl serverless <action> --help`.\n\n### Templates (alias: tpl)\n\n```bash\nrunpodctl template search <q>                        # find (+ template list [--type official/community/user, --all, --limit])\nrunpodctl template get <template-id>                 # details (README, env, ports)\nrunpodctl template create --name \"x\" --image \"img\" [--serverless]\nrunpodctl template delete <template-id>\n```\n\n### Network Volumes (alias: nv)\n\n```bash\nrunpodctl network-volume list                         # List all volumes\nrunpodctl network-volume get <volume-id>              # Get volume details\nrunpodctl network-volume create --name \"x\" --size 100 --data-center-id \"US-GA-1\"  # Create volume\nrunpodctl network-volume update <volume-id> --name \"new\"  # Update volume\nrunpodctl network-volume delete <volume-id>           # Delete volume\n```\n\nFor exact network volume flags, run `runpodctl network-volume <action> --help`.\n\n> **No storage-tier flag.** `create` provisions the data center's **default** tier — there's\n> no `--type`. To get a **High-Performance** volume, use the console (a ⚡ data center's toggle)\n> or a raw **v2 REST** call (`POST https://v2-rest.runpod.io/v2/network-volumes` with\n> `\"type\":\"HIGH_PERFORMANCE\"`) — or the MCP `create-network-volume` tool, which takes\n> `volumeType` (`STANDARD` | `HIGH_PERFORMANCE`). Tier is immutable after creation. Launch details: golden path [21](../runpod/golden-paths/21-storage-tiers.md).\n\n### Models (Model Repository)\n\n`runpodctl model` manages the **Runpod Model Repository** — managed, versioned storage\nfor your **own** model artifacts (upload once, distributed to workers; not pinned to a\ndata center like a network volume). What it is, why/how, migrating off a baked-in model,\nand Model-Repo-vs-volume: **[reference/model-caching.md](reference/model-caching.md)**.\n\n```bash\nrunpodctl model list                                  # List your models\nrunpodctl model list --all                            # List all models (not just yours)\nrunpodctl model list --name \"llama\"                   # Filter by name\nrunpodctl model list --provider \"meta\"                # Filter by provider\nrunpodctl model add --name \"my-model\" --model-path ./model   # Upload a local model dir (multipart)\nrunpodctl model remove --name \"my-model\" --owner <owner>     # Remove a model\n```\n\n`model add` supports upload sessions, versioning, metadata, and private-source credentials — see live `runpodctl model add --help`.\n\n### Info & SSH\n\n```bash\nrunpodctl user                                       # account info + balance (alias: me)\nrunpodctl gpu list                                   # available GPUs + $/hr + per-DC stock (+ --include-unavailable)\nrunpodctl datacenter list                            # datacenters (alias: dc)\nrunpodctl ssh info <pod-id>                          # SSH connection details (command + key; NOT an interactive session)\n```\n\n**`gpu list` carries pricing and placement data** — `securePricePerHr` /\n`communityPricePerHr` (explicitly `null` when that cloud doesn't offer the GPU) and a\n`dataCenterAvailability[]` breakdown. Read the breakdown, not just top-level\n`stockStatus` (which is only the *best* status across DCs), when a create has to\nschedule in a specific DC — and pass `--include-unavailable`, since the default listing\nhides no-stock GPUs and can omit one that has stock only in the DC you want. The prices\nare **pod on-demand** rates. Shape, stock-value vocabulary and the `\"none\"` vs\nomitted-key sentinel:\n[reference/output-and-errors.md](reference/output-and-errors.md#gpu-pricing-and-per-data-center-availability).\n\n`ssh info` gives connection details, not a session — if interactive SSH isn't available, run `ssh user@host \"command\"`. **Registry auth, `billing` history, and SSH key management** (`ssh add-key`/`remove-key`) are in [reference/command-reference.md](reference/command-reference.md).\n\n### File Transfer\n\n```bash\nrunpodctl send <path>                                # prints a one-time code, then blocks until the receiver connects\nrunpodctl receive <code>                             # positional code (no --code flag)\n```\n\nEncrypted/incremental/compressed — don't pre-tar. **Key gotchas:** capture the **first line of `send` stdout** (the code) as it streams (background + tee), each `send` mints a **fresh** code, both sides must exit `0`. Full agent flow (pod push via `ssh` + `receive`): [reference/command-reference.md](reference/command-reference.md#file-transfer).\n\n### Utilities\n\n```bash\nrunpodctl doctor                                      # Diagnose and fix CLI issues\nrunpodctl update                                      # Update CLI\nrunpodctl version                                     # Show version\nrunpodctl completion                                  # Auto-detect shell and install completion\n```\n\n## URLs\n\n### Pod URLs\n\nAccess exposed ports on your pod:\n\n```\nhttps://<pod-id>-<port>.proxy.runpod.net\n```\n\nExample: `https://abc123xyz-8888.proxy.runpod.net`\n\n### Serverless URLs\n\n```\nhttps://api.runpod.ai/v2/<endpoint-id>/run        # Async request\nhttps://api.runpod.ai/v2/<endpoint-id>/runsync    # Sync request\nhttps://api.runpod.ai/v2/<endpoint-id>/health     # Health check\nhttps://api.runpod.ai/v2/<endpoint-id>/status/<job-id>  # Job status\n```\n\n`serverless create`/`get`/`list`/`update` already return `run`/`runsync`/`health` in a\n`urls` object — prefer those over hand-assembling, since a non-default\n`RUNPOD_INVOKE_URL` changes the base. Only `status/<job-id>` has to be built by hand.\n\n## Source & docs\n\n- CLI source: https://github.com/runpod/runpodctl\n- Releases (binaries): https://github.com/runpod/runpodctl/releases\n- Docs: https://docs.runpod.io/runpodctl/overview\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}