{"id":12770,"plugin_id":"Plugin_cd395a620ba481918f5e3b37ce9a123d","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:02:09.498Z","digest":"8e7da5085adb87370cd03c4e29d32d7cf6e317c9e37b2021ffdfa4ca2c693486","against":null,"payload":{"name":"flash","description":"runpod-flash — code-first serverless: write Python locally, run it on remote Runpod GPUs/CPUs with `flash dev` (hot-reload + live worker logs), then `flash deploy`. Use for @Endpoint/@remote functions, resource config, and debugging flash deployments. For CLI-only infra management use runpodctl or runpod-mcp.","included_files":[{"relative_path":"evals/client-external-image.eval.md","size_in_bytes":1012},{"relative_path":"evals/connect-existing-endpoint.eval.md","size_in_bytes":932},{"relative_path":"evals/cpu-gpu-pipeline.eval.md","size_in_bytes":1058},{"relative_path":"evals/dev-loop-iteration.eval.md","size_in_bytes":3484},{"relative_path":"evals/fixtures/dev-loop/main.py","size_in_bytes":786},{"relative_path":"evals/lb-multi-route-api.eval.md","size_in_bytes":1067},{"relative_path":"evals/qb-gpu-function.eval.md","size_in_bytes":966},{"relative_path":"reference/api.md","size_in_bytes":6196},{"relative_path":"reference/patterns.md","size_in_bytes":4837},{"relative_path":"reference/setup-and-cli.md","size_in_bytes":3209}],"skill_md_contents":"---\nname: flash\ndescription: >-\n  runpod-flash — code-first serverless: write Python locally, run it on remote\n  Runpod GPUs/CPUs with `flash dev` (hot-reload + live worker logs), then\n  `flash deploy`. Use for @Endpoint/@remote functions, resource config, and\n  debugging flash deployments. For CLI-only infra management use runpodctl or\n  runpod-mcp.\nuser-invocable: true\nmetadata:\n  author: runpod\n  version: \"1.1.2\" # x-release-please-version\nlicense: Apache-2.0\n---\n\n# Runpod Flash\n\nWrite code locally, iterate with `flash dev` — it runs your functions on remote Runpod GPUs/CPUs with hot-reload and live worker logs — then `flash deploy` to ship. `Endpoint` handles provisioning.\n\n**Load on demand — this skill keeps the mental model + gotchas inline; details live in [`reference/`](reference/):**\n\n| Need | Read |\n|------|------|\n| Install, auth, `flash init`, and the full `flash` command list | [reference/setup-and-cli.md](reference/setup-and-cli.md) |\n| `Endpoint(...)` constructor params, `NetworkVolume`/`PodTemplate`/`EndpointJob`, GPU & CPU enum tables | [reference/api.md](reference/api.md) |\n| Worked patterns — choosing a model, warm-worker model loading, CPU→GPU pipeline, parallel calls | [reference/patterns.md](reference/patterns.md) |\n\nQuick start: `uv tool install runpod-flash` → `flash login` (or `export RUNPOD_API_KEY=...`) → `flash init my-project` → `flash dev`. Details in [reference/setup-and-cli.md](reference/setup-and-cli.md).\n\n## Dev vs Deploy\n\n- `flash dev` — **iterate.** Local server at `:8888`, but your decorated functions\n  execute on **remote GPU/CPU workers**. Hot-reloads on save and **streams the worker's\n  logs live** to the terminal. No build/upload/deploy wait — use this the whole time you\n  develop.\n- `flash deploy` — **ship.** Builds an artifact and deploys a stable endpoint. Slow\n  (build + upload + provision); only do this once the code works under `flash dev`.\n\n`flash dev` ships **only the function body** to the worker, so a `NameError` for a\nmodule-level name surfaces immediately here. `flash deploy` imports the whole module and\ncan mask that bug (see Gotcha #1). Develop against `flash dev` and you catch it first.\n\n## Autonomous Dev Loop\n\n`flash dev` is a long-running server. Three rules:\n- **Run it in the background** — don't block on it.\n- **Capture its output** to a log file.\n- **Drive it over HTTP.**\n\nThe captured log is the remote worker's live stream (cold start, model load, `print`s,\ntracebacks) — read it to debug.\n\n```bash\nflash dev > /tmp/flash-dev.log 2>&1 &                          # background; never run it blocking\nfor i in $(seq 1 60); do grep -q \"flash dev  localhost:\" /tmp/flash-dev.log && break; sleep 2; done  # bounded ~2min; if it never appears, check the log for errors\nURL=$(grep -o \"localhost:[0-9]*\" /tmp/flash-dev.log | head -1)               # actual port (8888 bumps if taken)\ncurl -s \"$URL/main/predict\" -d '{\"data\": {...}}'               # dispatches to the remote worker\n```\n\n- **Read the real URL from the log** — flash auto-bumps the port if 8888 is in use, and\n  prints `✓ flash dev  localhost:<port>` plus the route table.\n- **Routes are namespaced by file**: `main.py`'s `/predict` is served at `/main/predict`.\n- **Two route shapes, two body shapes** (mismatch → `422` naming the missing field in `loc`):\n  - **Load-balanced** (`@api.post(\"/predict\")`) → `POST /main/predict`, body is the arg\n    at top level: a handler `def predict(data: dict)` wants `{\"data\": {...}}` (not the bare object).\n  - **Queue-based** (bare `@Endpoint` decorator) → `POST /main/runsync` (the local dev\n    server only generates `/runsync`; production also exposes `/run`),\n    body is **double-wrapped** in `input`: a handler `def synthesize(data: dict)` wants\n    `{\"input\": {\"data\": {...}}}`. The outer `input` is the queue envelope; the inner key is\n    the handler's param name.\n- Edit a handler and save — hot-reload re-syncs the body; just re-send the request, no\n  redeploy. Add `--auto-provision` to skip the first-call cold start. `kill %1` when done.\n\n## Endpoint: Three Modes\n\nFull constructor params and the GPU/CPU enum tables are in [reference/api.md](reference/api.md).\n\n### Mode 1: Your Code (Queue-Based Decorator)\n\nOne function = one endpoint with its own workers.\n\n```python\nfrom runpod_flash import Endpoint, GpuGroup\n\n@Endpoint(name=\"my-worker\", gpu=GpuGroup.AMPERE_80, workers=5, dependencies=[\"torch\"])\nasync def compute(data):\n    import torch  # MUST import inside function (cloudpickle)\n    return {\"sum\": torch.tensor(data, device=\"cuda\").sum().item()}\n\nresult = await compute([1, 2, 3])\n```\n\n### Mode 2: Your Code (Load-Balanced Routes)\n\nMultiple HTTP routes share one pool of workers.\n\n```python\nfrom runpod_flash import Endpoint, GpuGroup\n\napi = Endpoint(name=\"my-api\", gpu=GpuGroup.ADA_24, workers=(1, 5), dependencies=[\"torch\"])\n\n@api.post(\"/predict\")\nasync def predict(data: list[float]):\n    import torch\n    return {\"result\": torch.tensor(data, device=\"cuda\").sum().item()}\n\n@api.get(\"/health\")\nasync def health():\n    return {\"status\": \"ok\"}\n```\n\n### Mode 3: External Image (Client)\n\nDeploy a pre-built Docker image and call it via HTTP.\n\n```python\nfrom runpod_flash import Endpoint, GpuGroup, PodTemplate\n\nserver = Endpoint(\n    name=\"my-server\",\n    image=\"my-org/my-image:latest\",\n    gpu=GpuGroup.AMPERE_80,\n    workers=1,\n    env={\"HF_TOKEN\": \"xxx\"},\n    template=PodTemplate(containerDiskInGb=100),\n)\n\n# LB-style\nresult = await server.post(\"/v1/completions\", {\"prompt\": \"hello\"})\nmodels = await server.get(\"/v1/models\")\n\n# QB-style\njob = await server.run({\"prompt\": \"hello\"})        # optional: webhook=\"https://...\" for completion callback\nawait job.wait()\nprint(job.output)\n```\n\nConnect to an existing endpoint by ID (no provisioning):\n\n```python\nep = Endpoint(id=\"abc123\")\njob = await ep.runsync({\"prompt\": \"hello\"})  # runsync wraps this as {\"input\": {\"prompt\": \"hello\"}}\nprint(job.output)\n```\n\n## How Mode Is Determined\n\n| Parameters | Mode |\n|-----------|------|\n| `name=` only | Decorator (your code) |\n| `image=` set | Client (deploys image, then HTTP calls) |\n| `id=` set | Client (connects to existing, no provisioning) |\n\nThe table above is *how* the mode is picked from params. *When* to reach for `image=`:\n\n### When to use `image=` (custom container) vs your own code\n\nDefault to writing Python (decorator / routes) — it runs arbitrary code with\n`dependencies=[...]`/`system_dependencies=[...]` and needs no Dockerfile. Even large\nHuggingFace models stay in decorator mode (weights stream at runtime — see\n[reference/patterns.md → Loading ML models](reference/patterns.md#loading-ml-models-warm-workers)).\nReach for `image=` **only** when you need:\n\n- **a pre-built inference server** — vLLM, TensorRT-LLM (`image=\"vllm/vllm-openai:latest\"`, or `runpod/worker-vllm`, `runpod/worker-comfy`)\n- **system-level deps not pip-installable** — a specific CUDA/cuDNN, OS libraries\n- **models baked into the image** — to skip the runtime download entirely\n- **an existing Runpod Serverless worker** — you already have a working image\n\nTrade-off: `image=` mode **can't run arbitrary Python** (the image owns all logic) and the\nimage must implement a Runpod Serverless handler. Full list + examples:\nhttps://docs.runpod.io/flash/custom-docker-images\n\n## Gotchas\n\n1. **Only the function body ships to the worker** -- most common error. Put imports *and* any module-level constants/helpers the function uses *inside* the decorated body. `flash deploy` imports the whole module so module globals happen to work; `flash dev` ships just the body, so a module-level name raises `NameError`. A handler that works deployed can break under dev — fix it by moving everything inside.\n2. **Forgetting await** -- all decorated functions and client methods need `await`.\n3. **Missing dependencies** -- must list in `dependencies=[]`.\n4. **gpu/cpu are exclusive** -- pick one per Endpoint.\n5. **idle_timeout is seconds** -- default 60s, not minutes.\n6. **10MB payload limit** -- pass URLs, not large objects. Return binary (audio/images/files) as base64 in the JSON (`{\"audio_b64\": ...}`) and decode client-side; for larger outputs write to a NetworkVolume or upload to storage and return a URL.\n7. **Client vs decorator** -- `image=`/`id=` = client. Otherwise = decorator.\n8. **Auto GPU switching requires workers >= 5** -- pass a list of GPU types (e.g. `gpu=[GpuGroup.ADA_24, GpuGroup.AMPERE_80]`) and set `workers=5` or higher. The platform only auto-switches GPU types based on supply when max workers is at least 5.\n9. **`runsync` timeout is 60s** -- cold starts can exceed 60s. Use `ep.runsync(data, timeout=120)` for first requests or use `ep.run()` + `job.wait()` instead.\n10. **Request body shape (raw/external HTTP callers only)** -- match the request shape to the endpoint type:\n    - **LB routes** (`@api.post(...)`): send the handler arg at the top level — `{\"data\": {...}}`.\n    - **QB endpoints** (bare `@Endpoint`, hit via `.../run` or `.../runsync`): the worker calls\n      **`handler(**job_input)`**, so the request's `input` keys must match the handler's parameter\n      names — `def transcribe(input_data: dict)` wants `{\"input\": {\"input_data\": {...}}}`, and\n      `def read(input: dict)` wants `{\"input\": {\"input\": {...}}}`. A mismatch fails with\n      `got an unexpected keyword argument …`. Use `**kwargs` if the handler ignores the payload.\n    - **Never send an empty `input`.** A QB request with `{\"input\": {}}` is rejected by the\n      worker SDK as `Job has missing field(s): id or input` — always include at least one key.\n    - *Context:* the flash client (`ep.runsync(x)`, `api.post(...)`) hides the spreading, so this\n      only bites raw HTTP/external callers (mismatch behavior verified 2026-07-10 via worker logs).\n      See *Autonomous Dev Loop*.\n11. **Load a model once per worker (not per call)** -- for real inference use a class `@Endpoint` whose `__init__` loads the model once per worker (see [reference/patterns.md → Loading ML models](reference/patterns.md#loading-ml-models-warm-workers)). In function-form, reconcile with #1 by caching in a module global *inside* the body so it works under both `flash dev` and `deploy`:\n    ```python\n    global _MODEL\n    try: _MODEL\n    except NameError: _MODEL = load_model()   # runs once per worker, reused across calls\n    ```\n12. **Native CUDA libs go in `dependencies=[]` too** -- e.g. CTranslate2/faster-whisper needs `nvidia-cublas-cu12` + `nvidia-cudnn-cu12` or it silently falls back to CPU. Add them alongside the Python package.\n13. **Silent 401 auth failure** -- a set `RUNPOD_API_KEY` env var overrides the `flash login` token, so a bad/expired key wins. The failure is quiet: provisioning logs `GraphQL request failed: 401`, but `flash dev` still prints its normal ready line (\"failed endpoints deploy on-demand\"), so it *looks* healthy. When endpoints fail to provision:\n    1. Check the provisioning log for `GraphQL request failed: 401`.\n    2. Verify the current key independently: `curl -s -o /dev/null -w '%{http_code}' https://rest.runpod.io/v1/endpoints -H \"Authorization: Bearer $RUNPOD_API_KEY\"` (200 = good, 401 = bad).\n    3. Fix it: `unset RUNPOD_API_KEY` to fall back to the `flash login` token, or `export` a valid key.\n14. **`system_dependencies=` adds to cold start** -- apt packages (e.g. `[\"ffmpeg\", \"espeak-ng\"]`) install on the worker before first use, so the initial call is slower (on top of any model download); warm calls are unaffected.\n15. **Teardown a deployed app with `flash app delete <app>`** -- `flash undeploy list` may show \"no endpoints\" for an app that is deployed and serving; `flash app delete` (or `runpodctl serverless delete <id>`) reliably removes it.\n\n## Resources\n\n- Setup & CLI: [reference/setup-and-cli.md](reference/setup-and-cli.md) · API & compute enums: [reference/api.md](reference/api.md) · Patterns: [reference/patterns.md](reference/patterns.md)\n- Flash source: https://github.com/runpod/flash\n- Runnable examples: https://github.com/runpod/flash-examples — clone and adapt the closest one\n- Package (PyPI): https://pypi.org/project/runpod-flash/\n- Docs: https://docs.runpod.io/flash/overview\n  - Custom Docker images (when + how): https://docs.runpod.io/flash/custom-docker-images\n  - Storage / network volumes: https://docs.runpod.io/flash/configuration/storage\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}