← Files Runpod (Official)ARCHIVED FILE
skills/runpod/golden-paths/14-load-balancing-endpoint.md
15 KB · Oct 5, 2026 · 18:17 UTC
# Golden path 14 — load-balancing serverless endpoint (custom HTTP worker)
**Goal:** deploy a **load-balancing** Serverless endpoint — where your worker runs its own
HTTP server and Runpod routes requests **directly** to it (custom URL paths, single-hop, no
queue), as opposed to the queue/handler `/run` model — using a **custom image + headless
API** (no Console, no flash).
**Status:** ✅ COVERED — live-verified 2026-07-13. A stdlib-only FastAPI-free HTTP worker
(`<your-registry>/gp14-lb:v1`) was deployed as an `LB`-type endpoint via GraphQL `saveEndpoint`,
and `GET /ping`, `POST /echo`, `GET /stats` all returned `200` at
`https://<ENDPOINT_ID>.api.runpod.ai/<path>` with request state persisting on the worker.
**Lane(s):** custom Docker image + `runpodctl template create` + **GraphQL `saveEndpoint`
(`type: "LB"`)** or **v2 REST / Runpod MCP (`LOAD_BALANCER`)** + plain HTTP invocation.
(flash covers LB code-first; this is the image/API way.)
## When to use this
Reach for a load-balancing endpoint instead of the queue/handler model when you need:
- **Direct access to your model's own HTTP server** (vLLM's OpenAI-compatible server, a
Triton/TGI server, any FastAPI/Flask app) — expose it as-is, no handler wrapper.
- **Custom URL paths and HTTP verbs** (`POST /v1/chat/completions`, `GET /stats`,
WebSockets) rather than fixed `/run` + `/runsync`.
- **Lower, single-hop latency** for real-time apps and streaming.
- **Non-JSON payloads** or multiple logical endpoints inside one worker.
Stick with **queue-based** endpoints (golden paths 03/05/12) when you want guaranteed request
processing, automatic retries, and queue buffering under burst — the LB model **drops**
requests when overloaded (UDP-like) and has **no built-in retry** (queue-based is TCP-like).
| | Load balancing | Queue-based |
| --- | --- | --- |
| Request flow | direct to worker HTTP server | through the job queue |
| You implement | a full HTTP server (any framework) | a `handler(job)` function |
| API surface | your own paths/verbs | fixed `/run`, `/runsync`, `/status` |
| Under overload | drops requests, no retry | buffers in queue, auto-retries |
| Invoke URL | `https://<ID>.api.runpod.ai/<path>` | `https://api.runpod.ai/v2/<ID>/run` |
## The worker contract (the thing to get right)
A load-balancing worker is **just an HTTP server**, with two rules Runpod enforces:
1. **Expose a `GET /ping` health route.** The load balancer polls it and routes only to
workers that answer `200`:
| `/ping` returns | Worker state |
| --- | --- |
| `200` | healthy — receives traffic |
| `204` | still initializing (cold start) |
| anything else | unhealthy — pulled from the pool |
2. **Listen on the configured port.** Two env vars drive this — set **both explicitly** and
**expose that port** on the template (see the gotcha below):
| Env var | Meaning | Default |
| --- | --- | --- |
| `PORT` | main app server port | `80` |
| `PORT_HEALTH` | port the `/ping` probe hits | same as `PORT` |
Everything else — routes, request/response shape — is yours. Requests to
`https://<ENDPOINT_ID>.api.runpod.ai/<path>` land on your server's `<path>` unchanged; Runpod
enforces bearer auth (`Authorization: Bearer <RUNPOD_API_KEY>`) at the edge before routing.
## Prerequisites
- `RUNPOD_API_KEY` exported; Docker running; a Docker Hub (or other) registry login.
- `runpodctl` (any recent version — used only for `template create`).
- The endpoint type is set by **v2 REST** (`"type": "LOAD_BALANCER"` on `POST /v2/serverless`)
or by **GraphQL `saveEndpoint`** (`type: "LB"`). Neither the **v1** REST
`POST /v1/endpoints` body nor `runpodctl serverless create` exposes an endpoint-type field,
so those two can't produce a load balancer.
## Walkthrough (verified commands)
### 1. Build a tiny LB worker (an HTTP server, no Runpod SDK)
No `runpod` package, no handler — just a server that answers `/ping` plus your own routes.
This example uses the Python stdlib so the image is tiny and dependency-free:
```python
# app.py
import os, json, time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
PORT = int(os.getenv("PORT", "80")) # Runpod injects PORT; bind to it
START, COUNT = time.time(), 0
class H(BaseHTTPRequestHandler):
def _send(self, code, payload):
body = json.dumps(payload).encode()
self.send_response(code); self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body))); self.end_headers()
self.wfile.write(body)
def do_GET(self):
if self.path == "/ping": self._send(200, {"status": "healthy"}) # health probe
elif self.path == "/stats": self._send(200, {"requests": COUNT})
else: self._send(404, {"error": "not found"})
def do_POST(self):
global COUNT; COUNT += 1
n = int(self.headers.get("Content-Length", 0))
data = json.loads(self.rfile.read(n) or b"{}") if n else {}
if self.path == "/echo":
self._send(200, {"worker": os.getenv("RUNPOD_POD_ID", "?"),
"request_number": COUNT, "you_sent": data})
else: self._send(404, {"error": "not found"})
def log_message(self, *a): pass
ThreadingHTTPServer(("0.0.0.0", PORT), H).serve_forever()
```
```dockerfile
# Dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY app.py .
CMD ["python3", "app.py"]
```
(FastAPI/Flask work identically — the docs' reference worker uses FastAPI + uvicorn. The
contract is only "serve `/ping` on the port," not a specific framework.)
```bash
docker build --platform linux/amd64 -t <your-registry>/gp14-lb:v1 .
docker push <your-registry>/gp14-lb:v1
```
### 2. Create a serverless template — expose the port AND set PORT/PORT_HEALTH
This is the step that makes or breaks it (see [Gotchas](#gotchas)). Expose the chosen HTTP
port and set **both** `PORT` and `PORT_HEALTH` to it:
```bash
runpodctl template create --name gp14-lb-tmpl2 --serverless \
--image <your-registry>/gp14-lb:v1 --container-disk-in-gb 5 \
--ports "5000/http" --env '{"PORT":"5000","PORT_HEALTH":"5000"}'
# → template id, e.g. <template-id>
```
### 3. Create the endpoint as a load balancer
Two headless routes. Prefer **a** — v2 REST is the current shape on both the production
and dev hosts (verified 2026-07-29: both serve a byte-identical OpenAPI spec).
**a. v2 REST / Runpod MCP (preferred).** v2 REST takes `"type": "LOAD_BALANCER"` on
`POST /v2/serverless`. Load balancers have no queue, so the scaler must be
`REQUEST_COUNT` — the spec forbids `QUEUE_DELAY` for this type **on create**. Read the
invoke URL from the reply's `requestUrls.base` rather than assembling it:
```bash
curl -s -X POST https://v2-rest.runpod.io/v2/serverless \
-H "Authorization: Bearer $RUNPOD_API_KEY" -H 'Content-Type: application/json' \
-d '{"name":"gp14-lb-ep","image":"<your-registry>/gp14-lb:v1","type":"LOAD_BALANCER",
"gpu":{"pools":["ADA_24"],"count":1},"disk":5,"ports":["5000/http"],
"env":{"PORT":"5000","PORT_HEALTH":"5000"},
"workers":{"min":0,"max":1,"idleTimeout":5},
"scaling":{"type":"REQUEST_COUNT","requestCount":4}}'
# → {"id":"<endpoint-id>","type":"LOAD_BALANCER",
# "requestUrls":{"base":"https://<endpoint-id>.api.runpod.ai",
# "health":"https://<endpoint-id>.api.runpod.ai/ping"}, ...}
```
This is image-based (no template needed), so step 2 is optional on this route. `type` is
**required** on create (the spec lists it alongside `name`/`image`/`gpu`/`scaling`) and
**fixed** afterwards — `UpdateEndpointRequest` has no `type` field, so no PATCH can change
it.
> **Via the Runpod MCP server:** `create-endpoint` takes `endpointType: "LOAD_BALANCER"` and
> rejects `scalerType: QUEUE_DELAY` for it client-side, before spending a round trip — the
> shortest path when MCP is connected.
>
> If your server's `create-endpoint` has **no `endpointType` parameter**, it predates the v2
> serverless reshape and will 422 against production on any endpoint create or update. Upgrade
> it, or use the raw v2 REST call above meanwhile.
**b. GraphQL `saveEndpoint`.** The `type` field takes `QB`
(queue-based, the default) or `LB`. Use a browser `User-Agent` (Cloudflare) and pass the
api key in the query string:
```bash
curl -s -X POST "https://api.runpod.io/graphql?api_key=$RUNPOD_API_KEY" \
-H 'Content-Type: application/json' -H 'User-Agent: Mozilla/5.0' \
-d '{"query":"mutation($input:EndpointInput!){saveEndpoint(input:$input){id name type templateId}}",
"variables":{"input":{"name":"gp14-lb-ep","templateId":"<template-id>","type":"LB",
"gpuIds":"AMPERE_16","scalerType":"QUEUE_DELAY","scalerValue":4,
"workersMin":0,"workersMax":1,"idleTimeout":5}}}'
# → {"data":{"saveEndpoint":{"id":"<endpoint-id>","type":"LB",...}}}
```
`gpuIds` is required by `saveEndpoint` (tiers: `AMPERE_16`/`AMPERE_24`/`ADA_24`/`AMPERE_48`/
`ADA_48_PRO`/`AMPERE_80`/`ADA_80_PRO`). `workersMin: 0` = scale-to-zero. Confirm the type
stuck: the mutation echoes `"type": "LB"`.
GraphQL still accepts `scalerType: "QUEUE_DELAY"` on an `LB` endpoint (it is the legacy
flat scaler field and is not validated against the type here). It is not meaningful — a
load balancer has no queue to measure — and the v2 REST route above rejects the same
combination, so prefer `scalerType: "REQUEST_COUNT"` even on this route.
> **See also:** [17 — WebSocket worker](17-serverless-websocket.md) applies this same
> `type: "LB"` substrate to a WebSocket server — start here to understand the LB base path,
> then go there for a persistent-connection worker on top of it.
### 4. Warm the worker, then call your custom routes
Scale-to-zero means the first hit triggers a cold start. Poll the standard health API
(worker counts) until a worker is `ready`, then call your paths directly:
```bash
# trigger + wait for a healthy worker
curl -s "https://api.runpod.ai/v2/<endpoint-id>/health" -H "Authorization: Bearer $RUNPOD_API_KEY"
# {"workers":{"idle":1,"ready":1,"running":0,...}} ← ready:1 means routable
BASE="https://<endpoint-id>.api.runpod.ai" # ← the LB base URL: <ID>.api.runpod.ai
curl -s "$BASE/ping" -H "Authorization: Bearer $RUNPOD_API_KEY"
curl -s -X POST "$BASE/echo" -H "Authorization: Bearer $RUNPOD_API_KEY" \
-H 'Content-Type: application/json' -d '{"prompt":"golden path 14 lb","n":2}'
curl -s "$BASE/stats" -H "Authorization: Bearer $RUNPOD_API_KEY"
```
## Verify it works (observed 2026-07-13)
```text
GET /ping → 200 {"status": "healthy"}
POST /echo → 200 {"worker": "9agv3pjgc40qwb", "request_number": 1,
"you_sent": {"prompt": "golden path 14 lb", "n": 2}}
GET /stats → 200 {"requests": 1, "uptime_s": 22.4}
POST /echo → 200 {"worker": "9agv3pjgc40qwb", "request_number": 2, ...} # counter++, same worker
GET /ping (no Authorization header) → 401 # edge auth enforced
```
Two facts this proves about the LB model: requests hit **your** paths verbatim (there is no
`/run` indirection), and the in-memory `request_number` incremented `1 → 2` across calls —
you're talking to the worker's own long-lived process directly, not a stateless job.
## Gotchas
- **Set `PORT_HEALTH` and expose the port — a "running" worker is not a "ready" worker.**
The failure mode: the worker shows `running: 1` in `/health` but `ready: 0`, and every
request hangs until the LB's ~2-min "no worker available" timeout (`400 timed out waiting
for worker`, or a client-side timeout). That means the health probe never got a `200`.
The fix that worked here: expose the exact port on the template (`--ports "5000/http"`)
**and** set **both** `PORT` and `PORT_HEALTH` env vars to it. Relying on the documented
`PORT` default of `80` alone was not sufficient in practice — set them explicitly.
- **The type is set at creation, by v2 REST or GraphQL only.** v2 REST takes
`"type": "LOAD_BALANCER"`; GraphQL `saveEndpoint` takes `type: "LB"`. **v1** REST
`EndpointCreateInput` and `runpodctl serverless create` have no endpoint-type field. There is
no way to *flip* an existing endpoint either way — `UpdateEndpointRequest` has no `type`
field at all, so create it right or recreate it (or use the Console's **Endpoint Type →
Load Balancer** at creation). A queue-based worker image called on an LB path — or vice
versa — returns `{"error":"not allowed for QB API"}`.
- **Two URLs, don't mix them.** LB is invoked at `https://<ID>.api.runpod.ai/<path>`; the
queue API lives at `https://api.runpod.ai/v2/<ID>/run`. The `/health` worker-count endpoint
(`.../v2/<ID>/health`) still works for LB endpoints and is the cleanest readiness signal.
- **Cold starts need retries.** On scale-to-zero, expect a first-request miss while `/ping`
is still `204`. The docs recommend ≥3 retries with 5–10 s delays; here, polling `/health`
for `ready: 1` before sending real traffic was reliable.
- **No queue buffer, no retry.** Overload drops requests. If you need guaranteed processing,
use a queue-based endpoint instead.
- **Limits:** request timeout 2 min (no worker), processing timeout 5.5 min/request, payload
30 MB each way.
## Cost & cleanup
Endpoint is scale-to-zero (`workersMin: 0`), so idle cost is ~$0; a GPU worker only bills
during the brief warm test. Delete the endpoint (set workers to 0 first) and the template;
the image can stay in the registry.
```bash
# set workers to 0, then delete the endpoint
curl -s -X POST "https://api.runpod.io/graphql?api_key=$RUNPOD_API_KEY" \
-H 'Content-Type: application/json' -H 'User-Agent: Mozilla/5.0' \
-d '{"query":"mutation($i:EndpointInput!){saveEndpoint(input:$i){id workersMax}}",
"variables":{"i":{"id":"<endpoint-id>","name":"gp14-lb-ep","templateId":"<template-id>",
"type":"LB","gpuIds":"AMPERE_16","workersMin":0,"workersMax":0}}}'
curl -s -X POST "https://api.runpod.io/graphql?api_key=$RUNPOD_API_KEY" \
-H 'Content-Type: application/json' -H 'User-Agent: Mozilla/5.0' \
-d '{"query":"mutation{deleteEndpoint(id:\"<endpoint-id>\")}"}'
runpodctl template delete <template-id>
runpodctl serverless list # confirm the endpoint is gone
```
Kept image: `<your-registry>/gp14-lb:v1` (the tiny stdlib LB worker above).
## Skill gaps folded back
- The load-balancing endpoint type is settable headlessly two ways: **v2 REST**
(`"type": "LOAD_BALANCER"`) or **GraphQL `saveEndpoint`** (`type: "LB"`). **v1** REST
endpoint-create and `runpodctl serverless create` have no type field, and no API can change
the type after creation. Skills that create endpoints should note this when a custom-HTTP/LB
endpoint is wanted.
- **Setting `PORT_HEALTH` (and exposing the port) explicitly is effectively required**, not
optional — a worker that binds the app port but leaves `PORT_HEALTH` at its documented
default failed to become `ready`. Treat "expose the port + set `PORT` + set `PORT_HEALTH`"
as one atomic step for LB workers.
- Readiness for an LB endpoint is best observed via the queue-style `/v2/<ID>/health`
worker-count endpoint (`ready`/`running`/`initializing`) rather than by hammering `/ping`,
which blocks up to the 2-min no-worker timeout during cold start.
SHA-256: d93aa4199f621eebd9e4ebd1ddc80a3a3c65385e376f72d298cfb54a4dca884a