← Control PlaneCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Control Plane
Snapshot Sep 30, 2026 · 23:00 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "workload-troubleshooting",
"description": "Diagnoses unhealthy Control Plane workloads. Use when asked why a workload is crashing, not starting, OOMKilled, ImagePullBackOff, returning 502s, failing health checks, unreachable, or stuck deploying.",
"included_files": [],
"skill_md_contents": "---\nname: workload-troubleshooting\ndescription: \"Diagnoses unhealthy Control Plane workloads. Use when asked why a workload is crashing, not starting, OOMKilled, ImagePullBackOff, returning 502s, failing health checks, unreachable, or stuck deploying.\"\n---\n\n# Workload Troubleshooting\n\nThe symptom-first companion to the `workload` skill (which owns workload types, the spec shape, and the create/update tools). Given an unhealthy workload, map what you observe to its platform-specific root cause and a fix the schema will actually accept. Diagnosis is **read-only and MCP-first**; most failures trace to a Control Plane rule a generic engineer would not guess — deny-by-default firewalls, the secret identity+policy chain, blocked ports, the sleep-binary shutdown rule. The single most common is OOMKilled. Deep remediation for each area lives in the domain skill named in that section; this skill is the diagnostic map.\n\n## Step 1 — Gather state (read-only)\n\n| Tool | What it tells you |\n|---|---|\n| `list_deployments` | **Start here.** Per-location readiness with reason/message. Pass `location` to drill into one failing location. |\n| `get_workload_events` | Image pulls, crashes, scheduling, probe failures, `OOMKilled`. |\n| `get_workload_logs` | App logs (LogQL); the `_accesslog` container holds HTTP status codes and latency. |\n| `get_resource` (kind=`workload`) | The spec and current status. |\n| `list_metrics` then `query_metrics` | Resource pressure — memory before OOM, CPU, latency. |\n| `list_workload_replicas` | Confirm which replicas are currently running. Use the `cpln` CLI after reading the `cpln` skill if in-container inspection is essential. |\n| `query_traces` then `get_trace` | For a slow or intermittently failing request: which span in the path spent the time or errored. Needs tracing enabled on the GVC (opt-in). Deep dive in `metrics-observability`. |\n\nCLI fallback (MCP unavailable, interactive shell, or CI/CD):\n\n```bash\ncpln workload get WORKLOAD --gvc GVC -o json\ncpln workload eventlog WORKLOAD --gvc GVC -o json\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"}' --limit 50 # |= \"error\" filters; container=\"_accesslog\" for HTTP codes\ncpln workload connect WORKLOAD --gvc GVC --location LOCATION # interactive shell\n```\n\n## Failure catalog\n\n### Out of memory (OOMKilled) — the most common issue\n\n**Symptoms:** container restarts repeatedly, events show `OOMKilled`, crashes under load.\n\n`memory` is a hard cap — exceed it (app + runtime + GC + buffers) and the kernel kills the container. Usual culprits: Java without `-Xmx`, Node without `--max-old-space-size`, Python loading large datasets. Each container, sidecars included, has its own limit. With **Capacity AI** on, spiky workloads can be downsized too aggressively — set `minMemory` as a floor (Capacity AI never downscales CPU below 25 millicores).\n\n**Fix:** check real usage with `query_metrics`, then raise `memory` — but memory (MiB) must stay within **8× CPU (millicores)**, so raise `cpu` alongside it or the update is rejected (the default `cpu: 50m` caps memory at 400Mi). See *Apply fixes within the schema's limits*.\n\n### Image pull failures\n\n**Symptoms:** events show `ImagePullBackOff` / `ErrImagePull`, deployment stuck.\n\n- **Reference format** — `//image/NAME:TAG` (org registry), bare `NAME:TAG` (Docker Hub, no `docker.io/`), full URL (other registries).\n- **Platform** — images must be `linux/amd64` for managed locations (BYOK allows more).\n- **Pull secret** — a private external registry needs a pull secret in the GVC's `pullSecretLinks`; only `docker`, `ecr`, `gcp` secret types work, and org `//image/` images need none.\n\nThe registry secret must already exist (created by the user — offer a manifest scaffold for them to fill and apply, `setup-secret` skill); attach with `update_gvc`. Deep setup: `image` skill.\n\n### Secret access failures\n\n**Symptoms:** env vars empty, logs show missing config, secret-access errors in events.\n\nA workload reaches a secret only with all three in place: an **identity linked** to it (`spec.identityLink`), a **policy granting that identity `reveal`** on the secret, and a **correct reference**. Fastest fix: `grant_workload_secret_access` builds the whole chain in one call. Reference format by type:\n\n| Type | Reference |\n|---|---|\n| Opaque (decoded / raw) | `cpln://secret/NAME.payload` / `cpln://secret/NAME` |\n| Dictionary | `cpln://secret/NAME.KEY` |\n| Username & password | `cpln://secret/NAME.username` / `.password` |\n| Keypair | `cpln://secret/NAME.secretKey` / `.publicKey` / `.passphrase` |\n| TLS | `cpln://secret/NAME.cert` / `.key` / `.chain` |\n| AWS | `cpln://secret/NAME.accessKey` / `.secretKey` / `.roleArn` / `.externalId` |\n\nThe manual chain: `access-control` and `setup-secret` skills.\n\n### Port mismatch — healthy but 502/503\n\n- The spec port must match what the process listens on — compare the workload spec with application startup logs. If that is inconclusive, use the `cpln` CLI after reading the `cpln` skill for an in-container socket check.\n- On serverless, the runtime injects `PORT` and rejects a `PORT` env var that doesn't equal the exposed port.\n- **Type rules:** serverless exposes exactly one port (zero is rejected; TCP needs a dedicated LB — see *Dedicated load balancer & domain*); standard and stateful expose zero or more; cron serves no endpoint.\n- **Blocked ports** (cannot bind, invalid for TCP probes): `8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000`.\n\n### Firewall blocking traffic\n\n**Symptoms:** unreachable externally, can't reach external APIs, or can't talk to other workloads.\n\nDeny-by-default: external inbound disabled, external outbound disabled, internal `none`. Fix via `update_workload`:\n\n- **Inbound** — `external.inboundAllowCIDR` (e.g. `0.0.0.0/0`, or specific CIDRs).\n- **Outbound** — `external.outboundAllowCIDR`, or `outboundAllowHostname` (hostname rules reach only ports 80, 443, 445).\n- **Internal** — `internal.inboundAllowType`: `same-gvc` / `same-org` / `workload-list` (default `none` blocks all workload-to-workload traffic).\n\nFull model: `firewall-networking` skill.\n\n### Health-check failures\n\n**Symptoms:** events show probe failures, replicas unready, restarts.\n\nDefault probes: serverless gets readiness + liveness TCP on the container port; standard, stateful, and cron have none (cron strips them). Common fixes: raise `initialDelaySeconds` (0-600; default 10 readiness / 60 liveness) for slow starts; raise `periodSeconds` (1-600, default 10) or `timeoutSeconds` (1-600, default 1) for over-aggressive probes; ensure an HTTP path returns 200-399. **Readiness** failure removes the replica from the pool and pauses rollout; **liveness** failure restarts it. An autoscaled workload with no real readiness probe gets traffic before it's ready, causing 502s on scale-up — add an `httpGet` readiness probe (it needs a port; defaults to the container's). Probe tuning: `workload-security` skill.\n\n### Resource limits & Capacity AI\n\n**Symptoms:** won't schedule, throttled, or Capacity AI not adjusting.\n\n- CPU/memory must fit org quota; `maxScale` × per-replica resources is enforced at scheduling.\n- **Capacity AI** does not apply with CPU-utilization or multi-metric autoscaling, or on stateful workloads — use `minCpu` / `minMemory` instead (those persist; `capacityAI` is stripped on stateful).\n- **Stateful sizing** — `minCpu`/`cpu` at most 4000m apart (ratio ≤ 4:1); `minMemory`/`memory` at most 4096Mi apart (ratio ≤ 4:1).\n- **Ephemeral storage** — 1GB per CPU core (minimum 1GB); exceeding it replaces the replica.\n\nDeep model: `autoscaling-capacity` skill.\n\n### Container won't start (restrictions)\n\n- **UID 1337** is the mesh proxy's UID — running as it excludes the container from the sidecar, disabling mesh communication and mTLS. Override `runAsUser` to another UID in 1-65534 (0/root is rejected).\n- **Reserved container names** — not `istio-proxy` / `queue-proxy` / `istio-validation` (or other reserved names); cannot start with `cpln-` or `debugger-`.\n- **Reserved env vars** — names starting `CPLN_`, plus `K_SERVICE` / `K_CONFIGURATION` / `K_REVISION`, are rejected; each value caps at 4096 characters.\n- **Suspended** — `spec.defaultOptions.suspend: true` stops the workload (min/max scale 0). Clear it, or `cpln workload start WORKLOAD --gvc GVC`.\n\n### Autoscaling misconfiguration\n\n**Symptoms:** won't scale, 502s on scale-up, scale-to-zero not working, or an invalid-strategy error on create.\n\nStrategies: `concurrency`, `cpu`, `memory`, `rps`, `latency`, `keda`, `disabled`. Per type: **serverless** has no `latency` or multi-metric; **standard** has no `concurrency`; **stateful** has no `concurrency`; **cron** has no autoscaling at all (the block is removed). **Scale-to-zero:** serverless with `rps` or `concurrency`; standard and stateful only with `metric: keda` (otherwise the update is rejected); cron cannot. For 502s on scale-up, fix the readiness probe (above). Details: `autoscaling-capacity` skill.\n\n### Termination / graceful shutdown\n\n**Symptoms:** requests fail during deploys or scale-down, 502/503 on rollout, containers killed abruptly.\n\n- **Missing `sleep`** — if `sleep` is absent from **any** container, **all** containers get SIGKILL immediately (no drain). Many distroless/minimal images lack it; confirm with `which sleep`. Fix: include `sleep` or add a custom `preStop`.\n- **preStop error** — a failing custom `preStop` in any container SIGKILLs all of them.\n- **Ignores SIGTERM** — after the preStop (default `sleep 45`) the container gets the termination signal, then SIGKILL once `terminationGracePeriodSeconds` (default 90; max 900 without the `cpln/relaxGracePeriodMax` tag) elapses.\n\nSequence and rollout options: `workload-security` skill.\n\n### Volume mount failures\n\n**Symptoms:** can't read mounted files, permission denied, empty cloud volume.\n\n- **Secret volumes** need the identity + `reveal` policy chain (as *Secret access failures*).\n- **Cloud volumes** (S3, GCS, Azure Blob/Files) need an identity, a cloud-access policy, and outbound firewall to the provider hosts (`*.amazonaws.com`, `*.googleapis.com`, `*.blob.core.windows.net` / `*.file.core.windows.net` plus `*.azure.com`); auth is identity-only — embedded keys do not work. Read-only except Azure Files.\n- **Reserved mount paths**: `/dev`, `/dev/log`, `/tmp`, `/var`, `/var/log`. Max 15 volumes; no path may be a parent of another.\n- `filesystemGroupId` defaults to 0 (root) when unset — set it (1-65534) for a non-root app.\n\nVolume sets: `stateful-storage` skill.\n\n### Service-to-service failures\n\n**Symptoms:** a workload can't reach another internally (connection refused or timeout).\n\n- **Target's internal firewall** must allow the caller — `same-gvc` / `same-org` / `workload-list` (default `none`); listing a workload needs `view` on it.\n- **Endpoint** — `http://WORKLOAD.GVC.cpln.local:PORT` (use `http://`; the sidecar adds mTLS). An omitted port defaults to the target's first container port; only listed ports are reachable.\n- **Cross-GVC** — the target must allow `same-org` or list the caller; traffic may span locations and incur egress charges.\n\n### Dedicated load balancer & domain\n\n**Symptoms:** unreachable after enabling a dedicated LB, TCP broken, wrong Host header.\n\n- **TCP** needs a dedicated LB on the GVC plus a custom Domain with a TCP port — not on default endpoints (HTTP/HTTP2/gRPC only).\n- Enabling/disabling a dedicated LB causes brief DNS-propagation downtime and per-location charges.\n- **Serverless Host header** — a custom domain delivers the canonical endpoint as `Host` (the custom domain moves to `X-Forwarded-Host`); standard and stateful workloads get the custom domain as `Host`.\n- **Protocol compatibility** — the domain port protocol must match the container's: HTTP2 fronts HTTP2 or gRPC, HTTP fronts HTTP.\n\nRouting, TLS, and LBs: `domain` and `ipset-load-balancing` skills.\n\n## Apply fixes within the schema's limits\n\nA fix the Joi schema rejects at `update_workload` time is worse than none. Before applying a resource or option change, confirm it stays within these (the validator names the violated rule):\n\n- `memory` (MiB) at most 8× `cpu` (millicores); CPU ≥ 25m; memory ≥ 32MiB; `minCpu`/`minMemory` never above `cpu`/`memory`.\n- `runAsUser` / `filesystemGroupId`: 1-65534 (0/root rejected).\n- `terminationGracePeriodSeconds`: ≤ 900 (higher only with the `cpln/relaxGracePeriodMax` tag).\n- `capacityAI` with `metric: cpu` is rejected; `capacityAI` is stripped on stateful, cron, and vm.\n- Standard/stateful `minScale: 0` requires `metric: keda`; cron and vm cannot scale to zero.\n- A metric outside the workload type's allow-list is rejected.\n\nPrefer `update_workload` (PATCH — only the fields you set change). For full manifest control, author against `get_resource_schema` then `cpln apply -f workload.yaml --gvc GVC`.\n\n## Verify\n\nAfter applying, poll `list_deployments` until every location reports ready, confirm the original symptom cleared (events/logs), and report the **canonical endpoint** `list_deployments` returns — never a constructed URL. For a public workload, confirm it actually responds, not just that it is ready.\n\n## Troubleshooting\n\n| Symptom | Likely cause | First check |\n|---|---|---|\n| Restarts; `OOMKilled` in events | memory cap too low (or Capacity AI downsized) | `query_metrics` memory; raise `memory` + `cpu` |\n| `ImagePullBackOff` / stuck | bad image ref, wrong platform, missing pull secret | events; GVC `pullSecretLinks` |\n| Env vars empty | broken identity + `reveal` chain or wrong reference | `grant_workload_secret_access` |\n| Healthy but 502/503 | spec port ≠ listening port, or a blocked port | workload spec and startup logs |\n| Unreachable / can't call out | deny-by-default firewall | `firewallConfig` |\n| Won't become ready | probe path/port wrong or too aggressive | events; probe config |\n| Can't reach another workload | target internal firewall `none`, or `https://` used | target `inboundAllowType`; use `http://` |\n| 502/503 during deploys | missing `sleep`, or app ignores SIGTERM | `which sleep`; SIGTERM handling |\n| `update_workload` rejected | the fix violates a schema limit | *Apply fixes within the schema's limits* |\n\n## Quick reference\n\n### MCP tools\n\n| Tool | Purpose |\n|---|---|\n| `list_deployments` | Primary per-location readiness monitor |\n| `get_workload_events` | Image / crash / probe / schedule events |\n| `get_workload_logs` | App and `_accesslog` logs |\n| `list_workload_replicas` | List running replicas |\n| `query_metrics` (after `list_metrics`) | Memory / CPU / latency pressure |\n| `update_workload` | Apply a spec fix (PATCH) |\n| `grant_workload_secret_access` | Build the identity + `reveal` chain in one call |\n| `get_resource_schema` | Exact fields before a manifest-level fix |\n\n### CLI (fallback)\n\nUse when MCP is unavailable, for an interactive shell, or in CI/CD (service-account `CPLN_TOKEN`).\n\n```bash\ncpln workload get WORKLOAD --gvc GVC -o json\ncpln workload eventlog WORKLOAD --gvc GVC -o json\ncpln workload connect WORKLOAD --gvc GVC --location LOCATION\ncpln apply -f workload.yaml --gvc GVC\n```\n\n### Related skills\n\n- `workload` — primary skill (types, spec shape, tool division); start here.\n- `workload-security` — probe tuning, termination, `securityOptions`, direct LBs.\n- `firewall-networking` — inbound / outbound / internal rules.\n- `autoscaling-capacity` — scaling strategies and Capacity AI.\n- `access-control` / `setup-secret` — the identity + `reveal` chain.\n- `stateful-storage` — volume sets; `domain` / `ipset-load-balancing` — routing and LBs; `image` — registries and pull secrets.\n\n## Documentation\n\n- [Workload Types](https://docs.controlplane.com/reference/workload/types.md)\n- [Containers](https://docs.controlplane.com/reference/workload/containers.md)\n- [Capacity AI](https://docs.controlplane.com/reference/workload/capacity.md)\n- [Firewall](https://docs.controlplane.com/reference/workload/firewall.md)\n- [Termination](https://docs.controlplane.com/reference/workload/termination.md)\n"
}SHA-256: 6b97a3525bcf816feb9ffd4c64865b4782ccfdd0376cbafcb2c0c36ec58e8420