← Control PlaneCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Control Plane
Snapshot Sep 30, 2026 · 23:00 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "autoscaling-capacity",
"description": "Workload autoscaling and Capacity AI on Control Plane. Use when the user asks about scaling up/down, min/max replicas, scale-to-zero, concurrency/RPS/CPU/memory/latency scaling, KEDA, event-driven scaling, or right-sizing.",
"included_files": [],
"skill_md_contents": "---\nname: autoscaling-capacity\ndescription: \"Workload autoscaling and Capacity AI on Control Plane. Use when the user asks about scaling up/down, min/max replicas, scale-to-zero, concurrency/RPS/CPU/memory/latency scaling, KEDA, event-driven scaling, or right-sizing.\"\n---\n\n# Autoscaling & Capacity AI\n\nDeep skill for scaling and resource optimization. Everything scaling lives in **one block** — `spec.defaultOptions.autoscaling` (with `capacityAI` beside it); `spec.localOptions[]` overrides it per location. The platform keeps the chosen metric near but below `target`. For workload types, production defaults, and the spec shape, start with the **`workload`** skill.\n\n## Picking a metric\n\n| Metric | Scales on | Types | Notes |\n|---|---|---|---|\n| `concurrency` | avg in-flight requests per replica | **serverless only** (its default) | pair with `maxConcurrency` for a hard per-replica cap |\n| `rps` | requests per second per replica | all three | consistent-response-time HTTP |\n| `cpu` | % of allocated CPU | all three (standard/stateful default) | `target` ≤ 100; conflicts with Capacity AI (below) |\n| `memory` | % of allocated memory | all three | `target` ≤ 100 |\n| `latency` | response time in **ms** at `metricPercentile` | standard / stateful | `p50` (default) / `p75` / `p99`; `target` is ms, not % |\n| `multi[]` | several metrics; highest replica count wins | standard / stateful | entries from `cpu` / `memory` / `rps` only, each at most once; **replaces** `metric` and top-level `target` |\n| `keda` | external / event-driven triggers | standard / stateful | GVC must enable KEDA first; `target` is rejected |\n| `disabled` | nothing — fixed at `minScale` | all | realized as min = max |\n\nIf `metric` is omitted, serverless defaults to `concurrency`; standard/stateful default to `cpu`. A metric invalid for the workload type is **rejected** (e.g. `concurrency` on standard).\n\n**The metric constrains the type — decide them together.** Type is chosen at creation and is immutable, so a metric-type mismatch is a *type* problem, not a metric problem. The most common case: concurrency-style scaling on a standard workload — the fix is to create the workload as **serverless** (concurrency lives only there) or use **`rps`** on standard (the closest equivalent), not to retry with the same pairing.\n\n**Don't silently downgrade.** If a type constraint blocks the user's stated intent (concurrency scaling on stateful, Capacity AI on a CPU-scaled workload), surface the conflict with realistic alternatives and a recommendation — per the constraint-conflicts rule in the operating guide (`get_cpln_rules`). `disabled` with `min=max=1` is sometimes right (single-writer app), but say so explicitly.\n\n## The autoscaling block\n\nSet with `create_workload` / `update_workload`, then verify with `list_deployments`. All fields:\n\n```yaml\nspec:\n defaultOptions:\n autoscaling:\n metric: rps\n target: 100 # default 95; integer 1-20000; ≤100 for cpu/memory; ms for latency\n minScale: 2 # default 1; must be ≤ maxScale; 0 = scale-to-zero (rules below)\n maxScale: 10 # default 5; no schema maximum\n scaleToZeroDelay: 300 # 30-3600s, default 300\n maxConcurrency: 0 # serverless only; 0-30000, default 0 = unlimited (excess queues)\n metricPercentile: p99 # latency only: p50 (default) / p75 / p99\n capacityAI: true\n```\n\n- **Per-location overrides:** `spec.localOptions[]` (same fields + `location`) via `configure_workload_local_options` — also the only MCP home of `capacityAIUpdateMinutes`, `spot`, and `multiZone`; it replaces the full list.\n- **`scaleToZeroDelay` is dual-purpose:** on serverless it is the idle period before scaling to 0; on standard/stateful it sets the **scale-down stabilization window** (default 300s) — scale-up is immediate.\n\n### Multi-metric (standard/stateful)\n\n```yaml\nautoscaling:\n minScale: 2\n maxScale: 10\n multi:\n - metric: cpu\n target: 80\n - metric: memory\n target: 80\n```\n\nEach entry is evaluated independently; the highest replica count wins. Only `cpu` / `memory` / `rps`, each at most once; targets go inside the entries (`metric`/`target` at the top level are rejected alongside `multi`). With `multi`, Capacity AI defaults to off.\n\n## minScale / maxScale & scale-to-zero\n\n- **Production default is `minScale: 2`** for user-facing services; pick `1` only with a named reason (single-writer DB, leader election, dev/staging). `maxScale` stays at its default `5` unless the user names a maximum — set exactly what they name, never invent a cap.\n- **Scale-to-zero (`minScale: 0`) by type:** serverless — allowed freely; standard/stateful — **only with `metric: keda`** (anything else is rejected); cron — never. On serverless it reaches zero with `concurrency`/`rps`; `cpu`/`memory` ride an HPA that won't drop to zero.\n- **Never the AI's default** — even on serverless, even when the user said \"auto-scale\". Configure it only when the user asked for scale-to-zero by name; the next request after idle pays a cold start. Acceptable (still opt-in): rarely-used internal tools, dev/preview environments, KEDA workers behind a retry-tolerant queue. Full rule: the operating guide (`get_cpln_rules`).\n\n## KEDA (event-driven, standard/stateful)\n\n**1. Enable on the GVC first** — `update_gvc`:\n\n```yaml\nspec:\n keda:\n enabled: true # default false\n identityLink: //gvc/GVC/identity/NAME # optional: cloud/network access for the KEDA operator\n secrets: [//secret/NAME] # optional: each becomes a TriggerAuthentication named after the secret\n```\n\n**2. Set the workload** — `metric: keda` plus raw [KEDA trigger specs](https://keda.sh/) (passed through as-is):\n\n```yaml\nautoscaling:\n metric: keda # target is rejected with keda\n minScale: 0 # maps to KEDA minReplicaCount — this is how standard/stateful scale to zero\n maxScale: 10\n keda:\n triggers:\n - type: redis\n metadata:\n address: my-redis.my-gvc.cpln.local:6379\n queueLength: '5'\n passwordFromEnv: REDIS_PASSWORD\n```\n\n- Triggers needing auth reference a GVC-listed secret via `authenticationRef.name` (the TriggerAuthentication is named after the secret).\n- If the trigger source is a Control Plane workload, allow KEDA in the source's firewall: `internal.inboundAllowWorkload: [cpln://internal/keda]`.\n- Also supported: `keda.advanced.scalingModifiers` (custom formulas), `fallback`, `pollingInterval`, `cooldownPeriod`.\n- **Prometheus trigger** — scale on any platform or custom metric: `type: prometheus` with `serverAddress: https://metrics.cpln.io:443/metrics/org/ORG`, a `query` (PromQL), `threshold`, and `customHeaders: Authorization=Bearer SERVICE_ACCOUNT_TOKEN` (service account needs `readMetrics`). **Before wiring any trigger, confirm the signal resolves:** `list_metrics` for real names/labels, then `query_metrics` to run the PromQL — a never-resolving signal pins the workload at `minScale`. Custom app metrics come from the container `metrics` block (see **metrics-observability**).\n\n## Capacity AI\n\nRight-sizes each container's **reserved** resources (what you're billed for) from usage history, between the `minCpu`/`minMemory` floor and the `cpu`/`memory` ceiling. **On by default for serverless and standard; stripped on stateful and cron.**\n\n```yaml\nspec:\n containers:\n - name: app\n cpu: '1000m' # ceiling (and the fixed allocation when Capacity AI is off)\n memory: '1Gi' # ceiling\n minCpu: '100m' # floor\n minMemory: '256Mi' # floor\n defaultOptions:\n capacityAI: true\n```\n\n- **With `metric: cpu`:** explicitly enabling Capacity AI is **rejected** (dynamic CPU allocation fights CPU-based scaling); left unset with `cpu` or `multi`, it silently defaults to **off**.\n- **GPU containers reject Capacity AI.**\n- Adjustments land **in place** on standard when the cluster supports pod resize (no restart; otherwise a rolling update); on serverless they roll a new revision. Throttle frequency with `capacityAIUpdateMinutes` (min 2 — via `localOptions` or `cpln apply`; not on create/update tools).\n- Idle floor is **25m** CPU, rising with memory at **1 millicore per 3 MiB**. A just-changed workload pauses adjustments while history rebuilds — apps that reserve resources at startup may not benefit.\n\n### Resource bounds (all types)\n\n- Floors: CPU ≥ `25m`, memory ≥ `32Mi`; `minCpu ≤ cpu`, `minMemory ≤ memory`; `memory(MiB) / cpu(millicores) ≤ 8` (32 with tag `cpln/relaxMemoryToCpuRatio`).\n- **Without Capacity AI** (standard/serverless, explicit off): `cpu`/`memory` are the fixed allocation; `minCpu`/`minMemory` are ignored.\n- **Stateful** has no Capacity AI, but `minCpu`/`minMemory` still work: they become the static **reserved** request while `cpu`/`memory` stay the burst ceiling. Constraints: max/min ratio ≤ **4** AND gap ≤ **4000m** CPU / **4096Mi** memory.\n- **GPU:** `nvidia` model `t4` (quantity up to 4) or `a10g` (exactly 1); strict per-model CPU/memory minimums — fetch exact numbers with `get_resource_schema` (`kind: workload`).\n- **Cost:** billing follows reserved resources, so Capacity AI (or stateful `minCpu`) directly lowers cost.\n\n## Type × scaling matrix\n\n| | standard | serverless | stateful | cron |\n|---|---|---|---|---|\n| Metrics | cpu, memory, latency, rps, multi, keda, disabled | concurrency, cpu, memory, rps, disabled | same as standard | none — autoscaling stripped |\n| Capacity AI | default on | default on | stripped | stripped |\n| Scale to zero | keda only | yes (concurrency/rps) | keda only | no |\n| Resize without restart | yes | no (new revision) | — | — |\n\n## Troubleshooting\n\n| Symptom | Check |\n|---|---|\n| Not scaling up | Does the signal exist? `list_metrics` then `query_metrics`; check `maxScale`; check replica readiness via `list_deployments` |\n| Not scaling down | Standard/stateful stabilization window = `scaleToZeroDelay` (default 300s); check `minScale` |\n| Scale-to-zero not happening | Serverless needs `concurrency`/`rps`; standard/stateful need `metric: keda`; check `scaleToZeroDelay` |\n| KEDA not triggering | KEDA enabled on the GVC? Trigger auth secret listed in `gvc.spec.keda.secrets`? Source firewall allows `cpln://internal/keda`? |\n| Capacity AI not adjusting | Restrictions (cpu metric, stateful, GPU); recent spec change pauses it; `capacityAIUpdateMinutes` throttle |\n| Replicas stuck at `minScale` | The scaling metric never resolves — verify the PromQL/trigger returns data |\n\n## Quick reference — MCP tools\n\n| Tool | Purpose |\n|---|---|\n| `create_workload` / `update_workload` | The `autoscaling` block (incl. `multi`, `keda`) and `capacityAI` |\n| `configure_workload_local_options` | Per-location overrides; `capacityAIUpdateMinutes`, `spot`, `multiZone` |\n| `update_gvc` | Enable KEDA on the GVC (`keda.enabled`, `identityLink`, `secrets`) |\n| `list_deployments` | Replica counts and readiness per location |\n| `get_workload_events` | Scaling/scheduling events and errors |\n| `list_metrics` / `query_metrics` | Discover metric names/labels, then verify the scaling signal — never guess |\n\n**CLI fallback** (read the `cpln` skill first): `cpln apply -f manifest.yaml` for the full spec incl. `capacityAIUpdateMinutes`; primary interface in CI/CD (`CPLN_TOKEN` + `cpln apply --ready`).\n\n## Related skills\n\n| Need | Skill |\n|---|---|\n| Workload types, production defaults, spec shape — start here | `workload` |\n| Custom `metrics` block, built-in metrics, PromQL | `metrics-observability` |\n| Scaling-event and per-execution cron logs | `logql-observability` |\n| Stateful sizing and volume sets | `stateful-storage` |\n\n## Documentation\n\n- [Autoscaling Reference](https://docs.controlplane.com/reference/workload/autoscaling.md)\n- [Capacity AI Reference](https://docs.controlplane.com/reference/workload/capacity.md)\n- [Custom Metrics Reference](https://docs.controlplane.com/reference/workload/custom-metrics.md)\n- [Export Metrics Guide](https://docs.controlplane.com/guides/export-metrics.md)\n"
}SHA-256: 916c22c85ff241b3d8f7ca119a7f9f73e962b01448863e992cfd1b246116e811