← Control PlaneCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Control Plane
Snapshot Sep 30, 2026 · 23:00 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "workload-security",
"description": "Production hardening for Control Plane workloads. Use when asked about JWT/Envoy auth, security context (runAsUser), health probe tuning, direct load balancers, geo-location headers, or graceful shutdown / termination.",
"included_files": [],
"skill_md_contents": "---\nname: workload-security\ndescription: \"Production hardening for Control Plane workloads. Use when asked about JWT/Envoy auth, security context (runAsUser), health probe tuning, direct load balancers, geo-location headers, or graceful shutdown / termination.\"\n---\n\n# Workload Security & Production Hardening\n\nDeep-dive companion to the `workload` skill, which owns workload types, the spec shape, and the readiness-vs-liveness model. Everything below is production-hardening detail for an existing workload.\n\n**Where settings live.** Health probes go inline in `containers[]` via `create_workload` / `update_workload`. Every other block here — `sidecar.envoy`, `loadBalancer`, `securityOptions`, `rolloutOptions` — is set by its own `configure_workload_*` tool, a set-or-clear PATCH on that one field (`remove: true` clears it).\n\n## Health Probes\n\nDefine `readinessProbe` (gate traffic) and `livenessProbe` (restart on failure) as distinct probes in the container spec.\n\n**Defaults by workload type:**\n- **Serverless** — a TCP readiness probe on the listening port is injected by default (plus default startup and liveness probes). Adequate, but an `httpGet` against a real endpoint catches more failure modes (DB unreachable, dependency timeout, deadlock).\n- **Standard / Stateful** — **no probes by default**; add them explicitly for any production workload.\n- **Cron** — probes are stripped (ignored).\n\nNo HTTP healthcheck? Use `tcpSocket` on the listening port as a baseline — don't run a long-lived workload probe-less.\n\n### Probe schema\n\nEach probe takes exactly one of `exec` / `grpc` / `tcpSocket` / `httpGet`, plus these timing fields:\n\n| Field | Range | Default |\n|---|---|---|\n| `initialDelaySeconds` | 0-600 | 10 (readiness) / 60 (liveness) |\n| `periodSeconds` | 1-600 | 10 |\n| `timeoutSeconds` | 1-600 | 1 |\n| `successThreshold` | 1-20 | 1 |\n| `failureThreshold` | 1-20 | 3 |\n\n`httpGet` omitting `port` defaults to the first container port; `httpGet.scheme` defaults to `HTTP`. Keep liveness looser than readiness (e.g. `periodSeconds: 30`) — restarts are expensive.\n\n```yaml\ncontainers:\n - name: api\n image: //image/api:v1.0\n ports: [{ number: 8080, protocol: http }]\n readinessProbe:\n httpGet: { path: /healthz/ready, port: 8080 }\n initialDelaySeconds: 5\n failureThreshold: 3\n livenessProbe:\n httpGet: { path: /healthz/live, port: 8080 }\n initialDelaySeconds: 30\n periodSeconds: 30\n```\n\n## JWT Authentication\n\nJWTs are validated at the Envoy sidecar before requests reach the workload, via the `jwt_authn` HTTP filter under `spec.sidecar.envoy`. Set it per-workload with `configure_workload_sidecar`, or org-wide-per-GVC by putting the same `sidecar.envoy` on the GVC (`update_gvc`) — it then applies to every workload in that GVC.\n\nMust-know rules (the filter is strictly validated, not passthrough):\n- The filter `name`, `typed_config.\"@type\"`, and `priority` (0-100) must be exact — copy them from the example.\n- Each provider needs a matching `clusters[]` entry (`STRICT_DNS` + TLS transport socket) so Envoy can fetch the JWKS over HTTPS. `remote_jwks.http_uri.cluster` must equal that cluster's `name`.\n- `rules` are first-match-wins. A rule with no `requires` lets matching paths through **without** a token — use it to exempt health/metrics endpoints.\n- `claim_to_headers` forwards JWT claims into request headers, so the workload gets identity context without re-parsing the token.\n- Provider/cluster names starting with `cpln_` are UI-managed and restricted (no `async_fetch` / `retry_policy`, and `cache_duration` must equal `http_uri.timeout`). For hand-authored configs use a **non-`cpln_`** name for full Envoy flexibility.\n\n```yaml\nspec:\n sidecar:\n envoy:\n clusters:\n - name: auth0\n type: STRICT_DNS\n load_assignment:\n cluster_name: auth0\n endpoints:\n - lb_endpoints:\n - endpoint:\n address:\n socket_address: { address: YOUR_TENANT.auth0.com, port_value: 443 }\n transport_socket:\n name: envoy.transport_sockets.tls\n http:\n - name: envoy.filters.http.jwt_authn\n priority: 50\n typed_config:\n \"@type\": type.googleapis.com/envoy.extensions.filters.http.jwt_authn.v3.JwtAuthentication\n providers:\n auth0:\n issuer: https://YOUR_TENANT.auth0.com/\n audiences: [https://api.example.com]\n remote_jwks:\n http_uri: { uri: https://YOUR_TENANT.auth0.com/.well-known/jwks.json, cluster: auth0, timeout: 5s }\n cache_duration: 300s\n claim_to_headers:\n - { header_name: X-User-Sub, claim_name: sub }\n rules:\n - match: { prefix: /healthz } # public, no token required\n - match: { prefix: / }\n requires: { provider_name: auth0 } # everything else needs a valid JWT\n```\n\nFor any OIDC provider (Auth0, Firebase, Cognito, Okta), set `issuer` to its issuer URL, `remote_jwks.http_uri.uri` to its JWKS endpoint (usually `/.well-known/jwks.json`), and `audiences` to your client ID or API identifier.\n\n## Security Options\n\n`spec.securityOptions` (set with `configure_workload_security`):\n\n| Field | Range | Purpose |\n|---|---|---|\n| `runAsUser` | 1-65534 | UID for all container processes |\n| `filesystemGroupId` | 1-65534 | GID applied to mounted volumes |\n\nNeither has a schema default — unset means the image's user and root/GID 0 for volumes. Set `runAsUser` to a non-root UID for defense in depth; set `filesystemGroupId` so containers can share access to mounted volumes (e.g. volume sets). **Not valid for `type: vm`** (the guest OS owns its own security context).\n\n**Trap — never `runAsUser: 1337`.** That is the mesh proxy's UID. A container running as 1337 is excluded from the Envoy redirect, so it bypasses the mesh entirely — losing mTLS and firewall enforcement and getting *unfiltered* egress.\n\n### Workload permissions\n\nBeyond standard `view` / `edit` (implies `view`) / `create` / `delete` / `manage`, the workload kind adds three non-obvious permissions: `connect` (interactive shell into a replica) and `configureLoadBalancer` (toggle direct/dedicated LBs) — **neither implied by `edit`** — plus `exec`, which implies `exec.runCronWorkload` + `exec.stopReplica`. Full policy and principal setup: `access-control` skill.\n\n## Direct Load Balancers\n\n`spec.loadBalancer.direct` exposes workload ports through a per-location cloud LB. Unlike the shared LB, it supports custom TCP/UDP ports and needs no domain registration — use it for non-HTTP protocols or custom ports. You are responsible for your own TLS certificates. Set with `configure_workload_load_balancer`.\n\n- `enabled` (bool, required) — when `false`, the LB is stopped and accrues no charges.\n- `ports[]` — each entry: `externalPort` (22-32768, required), `protocol` (`TCP`/`UDP`, required), `containerPort` (80-65535, excludes reserved ports), `scheme` (`http`/`tcp`/`https`/`ws`/`wss`, optional, sets the UI link scheme; default `https`).\n- `ipSet` (optional) — link to an IP set for reserved static IPs.\n\n**Reserved container ports (rejected):** 8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000.\n\n```yaml\nspec:\n loadBalancer:\n direct:\n enabled: true\n ports:\n - { externalPort: 5432, protocol: TCP, scheme: tcp, containerPort: 5432 }\n```\n\nEach direct-LB location also gets a public Geo DNS address with latency-based routing, usable as a CNAME target for custom domains.\n\n### Geo location headers\n\n`spec.loadBalancer.geoLocation` adds MaxMind GeoLite2 data to inbound HTTP requests under the header names you set in `headers.{asn,city,country,region}` (each max 128 chars; `enabled` defaults `false`). When enabled, at least one header is required, names must be unique, and existing same-named headers are replaced. HTTP-exposed workloads only. Pair with header-based firewall rules for geo-filtering (`firewall-networking` skill).\n\n### Replica-direct endpoints\n\n`spec.loadBalancer.replicaDirect: true` gives each replica its own stable hostname (default `false`, **stateful workloads only**):\n- External: `WORKLOAD-GVCALIAS-INDEX.LOCATION.controlplane.us`\n- Internal: `replica-INDEX.WORKLOAD.LOCATION.GVC.cpln.local:PORT`\n\n## Graceful Termination\n\n`spec.rolloutOptions` (set with `configure_workload_rollout`) governs how replicas are removed during scaling, version updates, Capacity AI rollouts, and maintenance.\n\n| Field | Range / values | Default |\n|---|---|---|\n| `minReadySeconds` | integer ≥0 | 0 |\n| `maxSurgeReplicas` | integer or percent | unset (not for `vm`) |\n| `maxUnavailableReplicas` | integer or percent | unset (not for `stateful`) |\n| `scalingPolicy` | `OrderedReady` / `Parallel` | `OrderedReady` (not for `vm`) |\n| `terminationGracePeriodSeconds` | 0-900 (3600 with the `cpln/relaxGracePeriodMax` tag) | 90 |\n\n`terminationGracePeriodSeconds` is the total budget for a replica to shut down before SIGKILL; raise it for long-running requests. Termination sequence:\n\n1. The load balancer removes the replica from the pool (up to ~10s) so new requests route elsewhere.\n2. The Control Plane sidecar drains in parallel: it holds for `grace - 10` seconds (default 80), then waits for in-flight connections to complete before shutting down.\n3. Containers run their `preStop` hook. With no custom hook, a default `sh -c \"sleep N\"` runs where `N` = half the grace period (45s at the 90s default), giving the LB time to stop routing.\n4. After `preStop`, the container receives **SIGINT** and has the remaining grace to exit, then **SIGKILL**. Handle the termination signal to drain cleanly.\n\n**Trap — immediate SIGKILL of *all* containers** if `sleep`/`sh` is missing in *any* container (common with distroless/minimal images) or a custom `preStop` errors in *any* container. A custom `preStop` must include a delay or connection check, and must be tested — an error there force-kills the whole replica.\n\n## Verify\n\n- After any change, poll `list_deployments` until each location reports ready — it surfaces probe failures per location.\n- `get_workload_events` gives the probe/liveness failure reason and message; `get_workload_logs` shows app-side errors.\n- For JWT, send a request with and without a valid token (expect 401 without) and confirm the `claim_to_headers` header arrives at the workload.\n\n## Troubleshooting\n\n| Symptom | Likely cause | Fix |\n|---|---|---|\n| Deploy never ready; events show probe failures | wrong `httpGet` path/port, or app slow to start | fix path/port; raise `initialDelaySeconds` / `failureThreshold` |\n| All requests 401 after adding JWT | no public-exemption rule, or `provider_name` mismatch | add a no-`requires` rule for health paths; match the rule's `provider_name` to a provider key |\n| JWT always rejected / JWKS fetch fails | cluster missing or `remote_jwks.http_uri.cluster` doesn't match a `clusters[].name` | add the `STRICT_DNS` + TLS cluster and align the names |\n| Replica SIGKILL'd instantly on rollout | `sleep`/`sh` missing, or custom `preStop` errors | use a `sleep`-capable image, or a native-sleep `preStop`; test the hook |\n| Direct LB port rejected | `containerPort` is reserved, or `externalPort` outside 22-32768 | pick a non-reserved container port; keep `externalPort` in range |\n| `securityOptions` / `replicaDirect` rejected | `type: vm` (securityOptions) or non-stateful (replicaDirect) | remove the field or change the workload type |\n\n## Quick reference\n\n### MCP tools\n\nAll `configure_workload_*` tools are `full`-profile, set-or-clear PATCH (`remove: true` clears).\n\n| Tool | Purpose |\n|---|---|\n| `configure_workload_sidecar` | Set/clear `spec.sidecar.envoy` — JWT / Envoy filter chain |\n| `configure_workload_load_balancer` | Set/clear `spec.loadBalancer` — direct LB, geo headers, replicaDirect |\n| `configure_workload_security` | Set/clear `spec.securityOptions` — `runAsUser`, `filesystemGroupId` |\n| `configure_workload_rollout` | Set/clear `spec.rolloutOptions` — termination grace, surge/unavailable |\n| `create_workload` / `update_workload` | Probes go inline in `containers[]` (update is PATCH, merges containers by name) |\n| `list_deployments` | Poll per-location readiness; surfaces probe failures |\n| `get_workload_events` | Probe / liveness failure reason + message |\n| `get_workload_logs` | App-side logs for security / probe issues |\n\n### CLI (fallback)\n\nUse the CLI when the MCP server is unavailable or unauthenticated, or in CI/CD (service-account `CPLN_TOKEN`).\n\n```bash\ncpln workload get WORKLOAD --gvc GVC -o yaml-slim > workload.yaml\n# edit, then apply (CI/CD: add --ready to block until deployed)\ncpln apply -f workload.yaml --gvc GVC\n```\n\n### Related skills\n\n- `workload` — primary skill (types, defaults, spec shape, tool division); start here.\n- `firewall-networking` — CIDR / header rules, geo-filtering, LB types.\n- `access-control` — policies, principals, the full permission model.\n- `autoscaling-capacity` — scaling and Capacity AI settings.\n- `stateful-storage` — volume sets (pair with `filesystemGroupId`).\n\n## Documentation\n\n- [Workload Security](https://docs.controlplane.com/reference/workload/security.md)\n- [JWT Auth](https://docs.controlplane.com/reference/workload/jwt-auth.md)\n- [Termination](https://docs.controlplane.com/reference/workload/termination.md)\n- [Load Balancing](https://docs.controlplane.com/reference/workload/load-balancing.md)\n"
}SHA-256: b4667cf48e4b30248e832280707dd981509e00fcec58c0ff68d0d9323dfc5d59