← Plugin catalog
Developer Tools
Control Plane
Control Plane Corporation v1.0.1
Publisher description
From the marketplace listing
Control Plane runs containerized apps across AWS, GCP, Azure, Oracle Cloud, and your own hardware as one platform. Build a GitHub or a GitLab repo into a container and deploy it to a live HTTPS URL, give it a custom domain, attach Postgres or Redis, and let it autoscale. Diagnose failures from logs, metrics, and events, and find the over-provisioned workloads inflating your cloud bill.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package31 files · 124 KBBrowse files →
Skill instructions
access-control10.2 KB
---
name: access-control
description: "Primary skill for access control, policies, and RBAC on Control Plane. Use when the user asks about permissions, policies, service accounts, user access, group membership, bindings, who can do what, least-privilege, or IAM."
---
# Access Control & Policies — Primary Skill
A **policy** targets one resource kind and binds **permissions** to **principals** (users, groups, service accounts, workload identities). The common failure is a policy that **exists but grants nothing** — a wrong `targetKind`, permission name, or principal link fails with no error — so read the policy back after writing.
## The model
| Layer | Scope | Controls |
|---|---|---|
| **Billing-account roles** | account-wide (set on the billing account, separate from org policies) | `billing_admin`, `billing_viewer`, `org_creator` — none grant org-resource access |
| **Org policies** | per resource kind | all day-to-day access |
A caller is allowed an action when some policy on the resource's kind has a binding that lists that action (or a permission implying it) **and** names the caller — or a group the caller is in.
## Policy anatomy
A policy = one **`targetKind`** + a **target scope** + **bindings**.
```yaml
kind: policy
name: app-secret-access
targetKind: secret # exactly one resource kind
targetLinks: # OR `target: all` OR `targetQuery:` (pick one)
- //secret/database-url
bindings: # ≤ 50
- permissions: [reveal, use]
principalLinks: # 1–200
- //gvc/production/identity/app-identity
```
**Limits:** `bindings` ≤ 50; `principalLinks` 1–200 per binding; `targetLinks` ≤ 200. `origin` is read-only — `default` (yours) or `builtin` (locked; see Built-ins).
**Target scope (pick one):** `target: all` (org-wide roles) · `targetLinks` (specific resources) · `targetQuery` (tag query — see **query-spec**).
**Valid target kinds:** the `targetKind` enum accepts all **33** resource kinds, but only the **20 kinds with permission schemas** (the ones `get_permissions` accepts) are meaningful targets: `workload, secret, gvc, identity, image, org, policy, group, serviceaccount, user, volumeset, domain, location, ipset, mk8s, cloudaccount, agent, auditctx, quota, task`.
**Create vs update:**
- `create_policy` builds **one** binding from `addPermissions` × (`addUsers`/`addGroups`/`addServiceAccounts`/`addIdentities`) — providing only one side is an **error** (a binding needs both). It also requires **exactly one** target scope (`targetAll` / `targetLinks` / `targetQuery`). For several distinct bindings, use `update_policy` `addBindings`.
- `update_policy` **merges** bindings (matched by exact permission set — you can't extend a set) but **replaces** targets (`targetLinks` wholesale, `removeTargetLinks` incremental, `targetAll` — mutually exclusive).
## Permissions
**`get_permissions` (`kind`) is the source of truth** — it returns the kind's exact permission list and its implication map. Confirm names with it before writing a policy; never hand-write them, and don't assume a kind only has `create` / `delete` / `edit` / `view` / `manage`.
Two traps that don't need a lookup:
- **`manage` implies every permission** for a kind — grant it only to true admins.
- **Secret values need `reveal`, not `read`** — the most common mistake.
Many kinds add non-obvious permissions beyond CRUD — e.g. secret `reveal`/`use`, image `pull`, workload `connect`/`exec.*`, serviceaccount `addKey`, user `invite`/`impersonate`, mk8s `clusterAdmin` — so pull the real set with `get_permissions`.
## Principals
| Type | Link |
|---|---|
| User | `//user/EMAIL` |
| Group (preferred) | `//group/NAME` |
| Service account | `//serviceaccount/NAME` |
| Workload identity | `//gvc/GVC/identity/NAME` (**GVC-scoped** — never `//identity/NAME`) |
An **identity never belongs to a group**, so authorize it only with a binding that names its exact link. At runtime the attached identity is the workload's API credential: calls from inside the workload with the injected `CPLN_TOKEN` against `CPLN_ENDPOINT` carry exactly the permissions policies grant that identity — nothing more.
### Groups
Members are **users and service accounts only** (≤ 200). Bind policies to groups, not individuals.
- **Create / edit:** `create_group` (`name`, `memberLinks`, `memberQuery`, `identityMatcher`); `edit_group` (`addMemberLinks` / `removeMemberLinks` — read first with `get_resource` (kind="group")).
- **Dynamic:** `memberQuery` matches users by tag query; `identityMatcher` matches identities by a `jmespath`/`javascript` expression.
### Service accounts (non-human / CI/CD)
- **Key:** `add_key_to_service_account` (`serviceAccountName`, **`keyDescription` required**, optional `groupName`) — **auto-creates the SA if missing** and returns the key **once** (save it; lost = revoke + remint). `create_service_account` makes one with no key.
- **Revoke:** `update_service_account` `removeKeys: [NAME]` (immediate). `delete_resource` (kind="service_account") revokes all keys.
- **CI/CD auth:** store the key as the `CPLN_TOKEN` secret/env var — the CLI uses it ahead of any profile (and works without one); don't pass `--token` on the command line (it leaks into logs). Full setup: **gitops-cicd**.
### Users (IDP-backed)
- **Invite:** `invite_user_to_org` (`email`, optional `groupName`).
- **Read / remove:** `get_resource` (kind="user") / `delete_resource` (kind="user") take **`identifier`** (id or email); `list_resources` (kind="user") has an `email` filter. No `create_user`/`update_user`.
## Built-ins (seeded per org)
| Resource | Name | Grants |
|---|---|---|
| Group | `superusers` | `manage` on every kind (creator auto-added) |
| Group | `viewers` | `view` on every kind |
| Service account | `controlplane` | platform-internal — off-limits |
| Policies | `superusers-KIND` / `viewers-KIND` | `origin: builtin`, `target: all` |
- **Grant org-admin by adding the principal to `superusers`** (or `viewers` for read-only) — don't recreate admin policies.
- Built-in **policies** can't be created/edited/deleted; built-in **groups** can't be deleted, but their membership is editable (you can't remove yourself from `superusers`).
- Any resource tagged **`cpln/protected=true` can't be deleted** until untagged.
## Common RBAC patterns
`create_policy` builds these; the `cpln apply -f` manifest is the policy-as-code equivalent for CI/CD.
- **Org admin** — add to `superusers`. Scoped admin — one policy per `targetKind` with `[manage]`.
- **GVC developer** — `workload` `[connect, create, delete, edit, exec, view]` + `secret` `[create, delete, edit, reveal, use, view]` on a `developers` group.
- **Read-only** — add to `viewers`.
- **CI/CD SA** — `workload` `[create, delete, edit, view]` + `image` `[create, pull, view]` + `secret` `[use, view]` on `//serviceaccount/cicd-deployer`.
- **Workload identity secret access** — the anatomy example above; full flow in **setup-secret** (create identity, bind, attach via `spec.identityLink`).
- **Auditor** — `policy` `[view]` + `auditctx` `[view]` on an `auditors` group.
## Standard flow
1. `get_cpln_rules` (once per mutating session) and read this skill.
2. Confirm the **org** (and **gvc** for identities) — never guess.
3. `get_permissions` for the kind — confirm permission names.
4. Read current state: `list_resources` (kind="policy") / `get_resource` (kind="policy") (+ `get_resource` kind="group" / kind="service_account").
5. Smallest change; if destructive, confirm blast radius.
6. `create_policy` / `update_policy` (+ group / SA / user tools).
7. **Verify:** read the policy back; confirm the bindings and target resolved.
## Quick reference — MCP tools
| Tool | Purpose | Key params |
|---|---|---|
| `get_permissions` | Permissions + implications for a kind | `kind` |
| `list_resources` (kind="policy") / `get_resource` (kind="policy") | List / read policies | `name` |
| `create_policy` | Create a policy | `name`, `targetKind`, `targetAll`/`targetLinks`/`targetQuery`, `addPermissions`, `addUsers`/`addGroups`/`addServiceAccounts`/`addIdentities` |
| `update_policy` | Update metadata, targets, bindings | `name`, `addBindings`/`removeBindings`, `targetLinks`/`removeTargetLinks`/`targetAll`, `targetQuery` |
| `delete_resource` (kind="policy") | Delete a policy (destructive) | `name` |
| `list_resources` (kind="group") / `get_resource` (kind="group") | List / read groups | `name` |
| `create_group` | Create a group | `name`, `memberLinks`, `memberQuery`, `identityMatcher` |
| `edit_group` | Add/remove members, update meta | `name`, `addMemberLinks`, `removeMemberLinks` |
| `delete_resource` (kind="group") | Delete a group (destructive) | `name` |
| `list_resources` (kind="service_account") / `get_resource` (kind="service_account") | List / read SAs (key metadata only) | `name` |
| `create_service_account` | Create an SA (no key) | `name`, `description` |
| `add_key_to_service_account` | Mint a key (auto-creates SA) | `serviceAccountName`, `keyDescription`, `groupName` |
| `update_service_account` | Update meta / revoke keys | `name`, `removeKeys` |
| `delete_resource` (kind="service_account") | Delete an SA (revokes all keys) | `name` |
| `list_resources` (kind="user") / `get_resource` (kind="user") | List / read users | `email` / `identifier` |
| `invite_user_to_org` | Invite a user by email | `email`, `groupName` |
| `delete_resource` (kind="user") | Remove a user (destructive) | `identifier` |
**CLI fallback** (read the `cpln` skill first; verify with `cpln <resource> --help`): policy-as-code in CI/CD (`CPLN_TOKEN` + `cpln apply -f`), `cpln RESOURCE permissions` to list a kind's permissions, `cpln policy access-report NAME` to audit a policy.
## Related skills
| Need | Skill |
|---|---|
| Query language for `targetQuery` / `memberQuery` | `query-spec` |
| Org creation, billing, profiles, SSO | `org-management` |
| Audit trail of policy / access changes | `audit-compliance` |
| Full workload secret-access flow | `setup-secret` |
| Credential-free cloud access (AWS / GCP / Azure / NGS) | `setup-cloud-access` |
| Workload identities, private-network connectivity | `native-networking` |
## Documentation
- [Access Control Concepts](https://docs.controlplane.com/concepts/access-control.md)
- [Policy Reference](https://docs.controlplane.com/reference/policy.md)
audit-compliance7.83 KB
---
name: audit-compliance
description: "Audit trail and compliance on Control Plane. Use when the user asks about audit logs, who changed what, change tracking, audit contexts, writing custom audit events, security monitoring, SOC 2, HIPAA, or PCI compliance."
---
# Audit Trail & Compliance
Every mutation on every Control Plane resource — via Console, CLI, API, Terraform, Pulumi, or MCP — is recorded automatically in an append-only, tamper-proof audit trail; nothing to configure. Most tasks are answering **"who changed what, when"** with `query_audit_events`. Custom audit contexts exist for one purpose only: letting your own workloads write their own audit events.
## The model
Events live in **audit contexts** (org-scoped namespaces):
| Context | Origin | Holds |
|---|---|---|
| `cpln` — built-in, one per org | `builtin` | every platform mutation, automatically |
| custom — user-created | `default` | only events your workloads POST to it |
- Platform events go **only** to `cpln`; custom contexts never receive them — querying one returns only what workloads wrote.
- The `origin` value `default` means "user-created", not "default for the org".
- **No audit context can be deleted** — there is no delete in the API, CLI, MCP, or Console, and `terraform destroy` only removes the context from Terraform state; events are append-only. The built-in `cpln` context can't be edited either. Creating a context is permanent.
## Event anatomy
Each event carries `id`, `eventTime` / `postedTime` / `receivedTime`, `requestId`, `context` (`org`, audit-context `name`, `location`, `gvcAlias`, `podId`, `remoteIp`, `eventSource`), `subject` (`name`, `email`), `resource` (`id`, `type`, `name`, `data` — the resource snapshot), `action.type` (`create` / `edit` / `delete` / `exec`), and `result` (`status`, `message`).
Secret snapshots are scrubbed: `resource.data` for a secret never includes the payload, so the audit trail itself stores no secret values.
## Querying events
### MCP — `query_audit_events`
```json
// all workload mutations in the last 24h
{ "kind": "workload", "org": "my-org", "since": "24h" }
// who changed a specific secret in the last 30 days
{ "kind": "secret", "name": "db-password", "since": "30d" }
// several policies in one call (merged, newest-first)
{ "kind": "policy", "names": ["admin", "readonly"], "since": "7d" }
// everything one subject touched across all secrets
{ "kind": "secret", "subject": "user@example.com", "since": "30d" }
// a custom context — kind matches the resource.type your workload wrote
{ "kind": "order", "context": "my-app-audit", "since": "7d" }
// a past window — relative from/to mean "that long ago from now"
{ "kind": "workload", "org": "my-org", "from": "3mo", "to": "1mo" }
```
Inputs: `kind` (required) · `name` **or** `names[]` (max 25, mutually exclusive; omit both for every resource of the kind) · `gvc` (**required** when `kind` is `workload` / `identity` / `dbcluster` / `volumeset` and a name is given) · `subject` (user email, full link, or bare service-account name) · `context` (default `cpln`) · `since` (default `7d`) **or** `from` / `to` (ISO 8601, or a relative duration meaning that long ago — units `m`, `h`, `d`, `w`, `mo`, `y`; months are `mo`, not `M`) · `limit` (default 50, max 1000). Results are merged, sorted newest-first, truncated to `limit`.
### CLI
Every resource type has an `audit` subcommand with the same flags:
```bash
cpln workload audit my-app --gvc my-gvc --org my-org --since 24h
cpln secret audit db-password --org my-org --subject user@example.com
cpln workload audit --org my-org --since 24h # omit the ref: every workload in the org
cpln workload audit my-app --gvc my-gvc \
--from 2025-10-23T07:00:00Z --to 2025-10-24T07:00:00Z
cpln workload audit my-app --gvc my-gvc --from now-3M --to now-1M # window: 3 months ago to 1 month ago
```
Flags: `--since` (default `7d`; mutually exclusive with `--from`/`--to`), `--from`/`--to` (ISO 8601, a relative duration like `7d` or `3M`, or `now-` prefixed like `now-30d` — the CLI accepts `M` for months; MCP only accepts `mo`), `--subject`, `--context` (default `cpln`), `--max` (default 50).
## Managing audit contexts
- **Create:** `create_audit_context` (`name`; `description` defaults to the name; `tags`). CLI: `cpln auditctx create --name my-app-audit --org my-org`.
- **Read:** `list_resources` (kind="auditctx") / `get_resource` (kind="auditctx").
- **Edit:** `edit_audit_context` — description and tags only; `origin` is immutable; the built-in `cpln` context rejects edits. CLI: `cpln auditctx update my-app-audit --set description="..."`.
- **Delete:** impossible by design (see The model) — don't promise it.
## Writing events from a workload
1. **Create a context** — `create_audit_context`.
2. **Create an identity** — `create_identity`.
3. **Grant `writeAudit`** — `create_policy` with `targetKind: auditctx`, the context in `targetLinks`, and a binding of `writeAudit` to the identity (binding shape: **access-control** skill).
4. **Attach the identity** to the workload via `spec.identityLink` — `update_workload`.
5. **POST from inside the container:**
```bash
curl -H "Content-Type: application/json" \
-X POST "http://127.0.0.1:43000/audit/org/${CPLN_ORG}/auditctx/my-app-audit?async=true" \
-d '{"resource": {"id": "order-1234", "type": "order"}, "action": {"type": "refund"}}'
```
- The sidecar serves `127.0.0.1:43000` in every workload pod and attaches the workload's identity automatically — no token or auth header needed. The write is rejected unless that identity has `writeAudit` on the context.
- Body: `resource.id` and `resource.type` are required; `eventTime` (ISO 8601, defaults to now), `subject`, `action.type`, and `result` are optional; unknown fields are rejected.
- `?async=true` is fire-and-forget; drop it to get the stored event's `id` back in the response.
## Permissions (kind `auditctx`)
`create`, `edit`, `view`, `manage`, and the two that matter: **`readAudit`** (query events) and **`writeAudit`** (post events) — both distinct from `view`, which only reads the context resource itself. `manage` implies all; `edit`, `readAudit`, and `writeAudit` each imply `view`. Confirm with `get_permissions` (`kind: auditctx`).
## Compliance
Control Plane is **PCI DSS Level 1** and **SOC 2 Type II** certified (audited by Prescient Assurance). For the SOC 2 report or PCI Attestation of Compliance, contact support on Slack or [support@controlplane.com](mailto:support@controlplane.com); the [PCI Responsibility Matrix](https://controlplane.com/downloads/Control_Plane_PCI_Responsibilities_Matrix.pdf) is public. Billing and payments are processed externally by Stripe — Control Plane stores no cardholder data.
## Quick reference — MCP tools
| Tool | Purpose | Key params |
|---|---|---|
| `query_audit_events` | Query events for a kind | `kind`, `name`/`names`, `gvc`, `subject`, `context`, `since`/`from`/`to`, `limit` |
| `create_audit_context` | Create a custom context (permanent) | `name`, `description`, `tags` |
| `list_resources` (kind="auditctx") / `get_resource` (kind="auditctx") | List / read contexts | `name` |
| `edit_audit_context` | Update description / tags | `name`, `description`, `tags`, `removeTagKeys` |
**CLI fallback** (read the `cpln` skill first; verify with `cpln auditctx --help`): `cpln RESOURCE audit [ref]`, `cpln auditctx create / get / update / query / access-report / permissions`. There is no delete.
## Related skills
| Need | Skill |
|---|---|
| Policy and binding shape for `writeAudit` / `readAudit` grants | `access-control` |
| Ship runtime logs to external destinations for retention | `external-logging` |
| CLI setup and command reference | `cpln` |
## Documentation
- [Audit Trail](https://docs.controlplane.com/core/audittrail.md)
- [Audit Context Reference](https://docs.controlplane.com/reference/auditctx.md)
- [Compliance](https://docs.controlplane.com/compliance.md)
autoscaling-capacity11.8 KB
---
name: autoscaling-capacity
description: "Workload autoscaling and Capacity AI on Control Plane. Use when the user asks about scaling up/down, min/max replicas, scale-to-zero, concurrency/RPS/CPU/memory/latency scaling, KEDA, event-driven scaling, or right-sizing."
---
# Autoscaling & Capacity AI
Deep skill for scaling and resource optimization. Everything scaling lives in **one block** — `spec.defaultOptions.autoscaling` (with `capacityAI` beside it); `spec.localOptions[]` overrides it per location. The platform keeps the chosen metric near but below `target`. For workload types, production defaults, and the spec shape, start with the **`workload`** skill.
## Picking a metric
| Metric | Scales on | Types | Notes |
|---|---|---|---|
| `concurrency` | avg in-flight requests per replica | **serverless only** (its default) | pair with `maxConcurrency` for a hard per-replica cap |
| `rps` | requests per second per replica | all three | consistent-response-time HTTP |
| `cpu` | % of allocated CPU | all three (standard/stateful default) | `target` ≤ 100; conflicts with Capacity AI (below) |
| `memory` | % of allocated memory | all three | `target` ≤ 100 |
| `latency` | response time in **ms** at `metricPercentile` | standard / stateful | `p50` (default) / `p75` / `p99`; `target` is ms, not % |
| `multi[]` | several metrics; highest replica count wins | standard / stateful | entries from `cpu` / `memory` / `rps` only, each at most once; **replaces** `metric` and top-level `target` |
| `keda` | external / event-driven triggers | standard / stateful | GVC must enable KEDA first; `target` is rejected |
| `disabled` | nothing — fixed at `minScale` | all | realized as min = max |
If `metric` is omitted, serverless defaults to `concurrency`; standard/stateful default to `cpu`. A metric invalid for the workload type is **rejected** (e.g. `concurrency` on standard).
**The metric constrains the type — decide them together.** Type is chosen at creation and is immutable, so a metric-type mismatch is a *type* problem, not a metric problem. The most common case: concurrency-style scaling on a standard workload — the fix is to create the workload as **serverless** (concurrency lives only there) or use **`rps`** on standard (the closest equivalent), not to retry with the same pairing.
**Don't silently downgrade.** If a type constraint blocks the user's stated intent (concurrency scaling on stateful, Capacity AI on a CPU-scaled workload), surface the conflict with realistic alternatives and a recommendation — per the constraint-conflicts rule in the operating guide (`get_cpln_rules`). `disabled` with `min=max=1` is sometimes right (single-writer app), but say so explicitly.
## The autoscaling block
Set with `create_workload` / `update_workload`, then verify with `list_deployments`. All fields:
```yaml
spec:
defaultOptions:
autoscaling:
metric: rps
target: 100 # default 95; integer 1-20000; ≤100 for cpu/memory; ms for latency
minScale: 2 # default 1; must be ≤ maxScale; 0 = scale-to-zero (rules below)
maxScale: 10 # default 5; no schema maximum
scaleToZeroDelay: 300 # 30-3600s, default 300
maxConcurrency: 0 # serverless only; 0-30000, default 0 = unlimited (excess queues)
metricPercentile: p99 # latency only: p50 (default) / p75 / p99
capacityAI: true
```
- **Per-location overrides:** `spec.localOptions[]` (same fields + `location`) via `configure_workload_local_options` — also the only MCP home of `capacityAIUpdateMinutes`, `spot`, and `multiZone`; it replaces the full list.
- **`scaleToZeroDelay` is dual-purpose:** on serverless it is the idle period before scaling to 0; on standard/stateful it sets the **scale-down stabilization window** (default 300s) — scale-up is immediate.
### Multi-metric (standard/stateful)
```yaml
autoscaling:
minScale: 2
maxScale: 10
multi:
- metric: cpu
target: 80
- metric: memory
target: 80
```
Each entry is evaluated independently; the highest replica count wins. Only `cpu` / `memory` / `rps`, each at most once; targets go inside the entries (`metric`/`target` at the top level are rejected alongside `multi`). With `multi`, Capacity AI defaults to off.
## minScale / maxScale & scale-to-zero
- **Production default is `minScale: 2`** for user-facing services; pick `1` only with a named reason (single-writer DB, leader election, dev/staging). `maxScale` stays at its default `5` unless the user names a maximum — set exactly what they name, never invent a cap.
- **Scale-to-zero (`minScale: 0`) by type:** serverless — allowed freely; standard/stateful — **only with `metric: keda`** (anything else is rejected); cron — never. On serverless it reaches zero with `concurrency`/`rps`; `cpu`/`memory` ride an HPA that won't drop to zero.
- **Never the AI's default** — even on serverless, even when the user said "auto-scale". Configure it only when the user asked for scale-to-zero by name; the next request after idle pays a cold start. Acceptable (still opt-in): rarely-used internal tools, dev/preview environments, KEDA workers behind a retry-tolerant queue. Full rule: the operating guide (`get_cpln_rules`).
## KEDA (event-driven, standard/stateful)
**1. Enable on the GVC first** — `update_gvc`:
```yaml
spec:
keda:
enabled: true # default false
identityLink: //gvc/GVC/identity/NAME # optional: cloud/network access for the KEDA operator
secrets: [//secret/NAME] # optional: each becomes a TriggerAuthentication named after the secret
```
**2. Set the workload** — `metric: keda` plus raw [KEDA trigger specs](https://keda.sh/) (passed through as-is):
```yaml
autoscaling:
metric: keda # target is rejected with keda
minScale: 0 # maps to KEDA minReplicaCount — this is how standard/stateful scale to zero
maxScale: 10
keda:
triggers:
- type: redis
metadata:
address: my-redis.my-gvc.cpln.local:6379
queueLength: '5'
passwordFromEnv: REDIS_PASSWORD
```
- Triggers needing auth reference a GVC-listed secret via `authenticationRef.name` (the TriggerAuthentication is named after the secret).
- If the trigger source is a Control Plane workload, allow KEDA in the source's firewall: `internal.inboundAllowWorkload: [cpln://internal/keda]`.
- Also supported: `keda.advanced.scalingModifiers` (custom formulas), `fallback`, `pollingInterval`, `cooldownPeriod`.
- **Prometheus trigger** — scale on any platform or custom metric: `type: prometheus` with `serverAddress: https://metrics.cpln.io:443/metrics/org/ORG`, a `query` (PromQL), `threshold`, and `customHeaders: Authorization=Bearer SERVICE_ACCOUNT_TOKEN` (service account needs `readMetrics`). **Before wiring any trigger, confirm the signal resolves:** `list_metrics` for real names/labels, then `query_metrics` to run the PromQL — a never-resolving signal pins the workload at `minScale`. Custom app metrics come from the container `metrics` block (see **metrics-observability**).
## Capacity AI
Right-sizes each container's **reserved** resources (what you're billed for) from usage history, between the `minCpu`/`minMemory` floor and the `cpu`/`memory` ceiling. **On by default for serverless and standard; stripped on stateful and cron.**
```yaml
spec:
containers:
- name: app
cpu: '1000m' # ceiling (and the fixed allocation when Capacity AI is off)
memory: '1Gi' # ceiling
minCpu: '100m' # floor
minMemory: '256Mi' # floor
defaultOptions:
capacityAI: true
```
- **With `metric: cpu`:** explicitly enabling Capacity AI is **rejected** (dynamic CPU allocation fights CPU-based scaling); left unset with `cpu` or `multi`, it silently defaults to **off**.
- **GPU containers reject Capacity AI.**
- Adjustments land **in place** on standard when the cluster supports pod resize (no restart; otherwise a rolling update); on serverless they roll a new revision. Throttle frequency with `capacityAIUpdateMinutes` (min 2 — via `localOptions` or `cpln apply`; not on create/update tools).
- Idle floor is **25m** CPU, rising with memory at **1 millicore per 3 MiB**. A just-changed workload pauses adjustments while history rebuilds — apps that reserve resources at startup may not benefit.
### Resource bounds (all types)
- Floors: CPU ≥ `25m`, memory ≥ `32Mi`; `minCpu ≤ cpu`, `minMemory ≤ memory`; `memory(MiB) / cpu(millicores) ≤ 8` (32 with tag `cpln/relaxMemoryToCpuRatio`).
- **Without Capacity AI** (standard/serverless, explicit off): `cpu`/`memory` are the fixed allocation; `minCpu`/`minMemory` are ignored.
- **Stateful** has no Capacity AI, but `minCpu`/`minMemory` still work: they become the static **reserved** request while `cpu`/`memory` stay the burst ceiling. Constraints: max/min ratio ≤ **4** AND gap ≤ **4000m** CPU / **4096Mi** memory.
- **GPU:** `nvidia` model `t4` (quantity up to 4) or `a10g` (exactly 1); strict per-model CPU/memory minimums — fetch exact numbers with `get_resource_schema` (`kind: workload`).
- **Cost:** billing follows reserved resources, so Capacity AI (or stateful `minCpu`) directly lowers cost.
## Type × scaling matrix
| | standard | serverless | stateful | cron |
|---|---|---|---|---|
| Metrics | cpu, memory, latency, rps, multi, keda, disabled | concurrency, cpu, memory, rps, disabled | same as standard | none — autoscaling stripped |
| Capacity AI | default on | default on | stripped | stripped |
| Scale to zero | keda only | yes (concurrency/rps) | keda only | no |
| Resize without restart | yes | no (new revision) | — | — |
## Troubleshooting
| Symptom | Check |
|---|---|
| Not scaling up | Does the signal exist? `list_metrics` then `query_metrics`; check `maxScale`; check replica readiness via `list_deployments` |
| Not scaling down | Standard/stateful stabilization window = `scaleToZeroDelay` (default 300s); check `minScale` |
| Scale-to-zero not happening | Serverless needs `concurrency`/`rps`; standard/stateful need `metric: keda`; check `scaleToZeroDelay` |
| KEDA not triggering | KEDA enabled on the GVC? Trigger auth secret listed in `gvc.spec.keda.secrets`? Source firewall allows `cpln://internal/keda`? |
| Capacity AI not adjusting | Restrictions (cpu metric, stateful, GPU); recent spec change pauses it; `capacityAIUpdateMinutes` throttle |
| Replicas stuck at `minScale` | The scaling metric never resolves — verify the PromQL/trigger returns data |
## Quick reference — MCP tools
| Tool | Purpose |
|---|---|
| `create_workload` / `update_workload` | The `autoscaling` block (incl. `multi`, `keda`) and `capacityAI` |
| `configure_workload_local_options` | Per-location overrides; `capacityAIUpdateMinutes`, `spot`, `multiZone` |
| `update_gvc` | Enable KEDA on the GVC (`keda.enabled`, `identityLink`, `secrets`) |
| `list_deployments` | Replica counts and readiness per location |
| `get_workload_events` | Scaling/scheduling events and errors |
| `list_metrics` / `query_metrics` | Discover metric names/labels, then verify the scaling signal — never guess |
**CLI fallback** (read the `cpln` skill first): `cpln apply -f manifest.yaml` for the full spec incl. `capacityAIUpdateMinutes`; primary interface in CI/CD (`CPLN_TOKEN` + `cpln apply --ready`).
## Related skills
| Need | Skill |
|---|---|
| Workload types, production defaults, spec shape — start here | `workload` |
| Custom `metrics` block, built-in metrics, PromQL | `metrics-observability` |
| Scaling-event and per-execution cron logs | `logql-observability` |
| Stateful sizing and volume sets | `stateful-storage` |
## Documentation
- [Autoscaling Reference](https://docs.controlplane.com/reference/workload/autoscaling.md)
- [Capacity AI Reference](https://docs.controlplane.com/reference/workload/capacity.md)
- [Custom Metrics Reference](https://docs.controlplane.com/reference/workload/custom-metrics.md)
- [Export Metrics Guide](https://docs.controlplane.com/guides/export-metrics.md)
cdn-rate-limiting10.1 KB
---
name: cdn-rate-limiting
description: "CDN caching and request rate limiting for Control Plane workloads. Use when the user asks about CDN, Cloudflare, CloudFront, edge caching, rate limiting, request throttling, per-key or per-route limits, or DDoS protection."
---
# CDN & Rate Limiting
Two edge concerns, both built from existing primitives — there is no CDN or rate-limit resource kind. A **CDN** is bring-your-own (Cloudflare / CloudFront) pointed at the workload's canonical endpoint; **rate limiting** is an Envoy ratelimit service you deploy, enabled per workload by `cpln/rateLimit*` **tags**. Assumes the `workload` primer (firewall deny-by-default, canonical URL rules, create-then-verify).
## CDN
The pattern: the CDN proxies your domain and uses the workload's **canonical endpoint** as origin — read it from `status.canonicalEndpoint` or `list_deployments`, never construct it. A **workload endpoint** gives precise geo-routing and per-workload failover; a **GVC endpoint** serves one CDN route for many domains or wildcard subdomains, but keeps sending traffic to every location even when that location's workload is down.
### Cloudflare
1. **DNS at Cloudflare:** proxied CNAME (orange cloud on) from your subdomain to the canonical endpoint; SSL/TLS mode **Full (strict)**.
2. **Origin certificate** (SSL/TLS, then Origin Server; RSA 2048) becomes a Control Plane **TLS secret** — created by the user with the cert and key (offer a manifest scaffold; `setup-secret` skill), **TLS chain left empty** (the origin cert is self-signed).
3. **Domain at Control Plane:** `create_domain` (CNAME DNS mode), `set_domain_tls` with the secret as the custom **server certificate**, `add_domain_route` to the workload. **The apex domain must be verified before configuring a subdomain.**
### Amazon CloudFront
1. **ACM public certificate in `us-east-1`** (CloudFront requires that region), DNS validation, covering your subdomain or wildcard.
2. **Distribution:** origin = the workload's canonical endpoint (BYOK: the per-location endpoint from `list_deployments`); alternate domain name = your subdomain; attach the ACM cert.
3. **DNS:** CNAME your subdomain to the distribution's `*.cloudfront.net` name.
### Lock out direct access
With a CDN in front, restrict the workload firewall so only CDN traffic reaches it: set `inboundAllowCIDR` to the provider's published ranges ([CloudFront IP list](https://d7uri8nf7uskq.cloudfront.net/tools/list-cloudfront-ips)) via `update_workload`, and keep the list current. BYOK locations must also admit the ranges in the cluster's ingress security group — the CloudFront list is large, so raise the VPC quota for rules per security group to at least **530**. Details: **firewall-networking**.
## Rate limiting
Tags on the target workload inject an [Envoy rate-limit filter](https://github.com/envoyproxy/ratelimit) into its **inbound sidecar**: each request makes a gRPC check (1s timeout) against a ratelimit service you deploy (Envoy ratelimit + Redis); over-limit requests get **HTTP 429**.
### 1. Deploy the ratelimit stack
One multi-resource manifest (no bundled-apply MCP tool — use the CLI):
```bash
cpln gvc create --name ratelimit --location LOCATION --org ORG # or create_gvc
cpln apply --file rate-limiting.yaml --org ORG --gvc ratelimit
```
The [example manifest](https://raw.githubusercontent.com/controlplane-com/examples/main/examples/rate-limiting/rate-limiting.yaml) creates the **ratelimit** workload (`envoyproxy/ratelimit`), a **redis** workload, the **ratelimit-config** opaque secret (the rules), and the identity + policy for secret access. It assumes the GVC is named `ratelimit` (edit it if yours differs) and pins an older `envoyproxy/ratelimit` image tag — substitute a newer tag if desired.
**As shipped, the manifest is a trial setup, not a production one:** both workloads run `minScale: 1` with `spot: true`. Because enforcement is fail-closed, the ratelimit stack is tier-1 infrastructure for every tagged workload — for production raise its `minScale` to 2+, set `spot: false`, and run the GVC in the same locations as the workloads it protects (every request pays the check's round trip). A Redis restart only resets counters; a ratelimit outage denies traffic.
### 2. Configure the rules
Have the user edit the `ratelimit-config` opaque secret, [Envoy ratelimit format](https://github.com/envoyproxy/ratelimit#configuration); `unit`: `second` / `minute` / `hour` / `day`:
```yaml
domain: cpln
descriptors:
- key: authorization
rate_limit:
unit: minute
requests_per_unit: 10
```
After editing, reload the config: `cpln workload force-redeployment ratelimit --gvc ratelimit`.
### 3. Tag the target workload
Set with `update_workload` (or at creation):
| Tag | Required | Default | Meaning |
|---|:-:|---|---|
| `cpln/rateLimitAddress` | **Yes** | — | Canonical endpoint **hostname** of the ratelimit workload (no scheme prefix) — nothing happens without it |
| `cpln/rateLimitDescriptors` | **Effectively yes** | `authority` | Comma-separated: `authorization`, `host`, `path` — the default matches none of them, so **no limiting is applied until you set this** |
| `cpln/rateLimitScheme` | No | `https` | `https` dials the service over TLS with SNI |
| `cpln/rateLimitPort` | No | `443` | Port of the ratelimit service |
| `cpln/rateLimitDomain` | No | `cpln` | Must match `domain` in the config secret |
| Descriptor | Buckets per | HTTP header |
|---|---|---|
| `authorization` | API key / token | `Authorization` |
| `host` | domain | `Host` |
| `path` | endpoint | `:path` |
**There is no per-IP descriptor** — these three are the only ones wired. For per-client-IP throttling use the CDN layer (e.g. Cloudflare rate-limiting rules); keep this stack for per-token and per-route limits. Don't invent a `remote_address` descriptor — it does nothing.
### Combining descriptors
`cpln/rateLimitDescriptors: authorization,path` produces **one compound key** (an ordered tuple), not two independent limits. The config must **nest in the same order** as the tag's comma order:
```yaml
domain: cpln
descriptors:
- key: authorization # first tag entry
descriptors:
- key: path # second tag entry — nested, not a sibling
rate_limit:
unit: minute
requests_per_unit: 60
```
This buckets per token-and-path **pair** (any values). Add `value:` under a key to pin one route or token (e.g. `key: path, value: /api/search`). Two genuinely independent limits are not expressible through the tags — the filter emits a single action tuple.
**Anonymous-bypass warning:** Envoy skips the rate-limit check entirely when a descriptor header is missing from the request (`skip_if_absent` defaults to false and the filter doesn't set it). With `authorization` in the list, requests **without** an `Authorization` header are never checked — including the sample 10/minute config above. Limit by `path` or `host` (always present) when anonymous traffic matters, or throttle it at the CDN.
### Traps (from the filter's actual wiring)
- **Fail-closed:** the filter is configured with `failure_mode_deny: true` and a 1s check timeout — if the ratelimit service is down, unreachable, or the address is wrong, **inbound requests are denied**, not passed through. Verify the ratelimit workload is Ready (`list_deployments`) **before** tagging the target, and treat a typo in `cpln/rateLimitAddress` as an outage.
- **Descriptors must be explicit:** the built-in default `authority` produces an empty action list — the filter runs but limits nothing.
- **Tags are not validated:** these are plain tags interpreted by the platform; a misspelled tag name silently does nothing.
- The address is resolved by DNS from inside the mesh — use the ratelimit workload's exact canonical endpoint hostname.
## Layering the full stack
Traffic passes, in order: the **CDN** (absorbs and caches), then the **firewall** (`inboundAllowCIDR` = CDN IPs only, so nobody bypasses the CDN), then **rate limiting** (catches abuse that passes the CDN), then the workload.
## Verify the setup
- **Limiting works:** send requests past the limit and expect `429` — with the sample 10/minute rule, the 11th call returns it: `for i in $(seq 1 11); do curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: test" https://CANONICAL_ENDPOINT/; done`. After any rules edit, `force-redeployment` the ratelimit workload first.
- **CDN serves:** `curl -I https://SUBDOMAIN` shows the CDN's header (`cf-cache-status` on Cloudflare, `x-cache` on CloudFront).
- **Bypass is closed:** `curl` the canonical endpoint directly — after the firewall lock-down it must no longer answer from outside the CDN ranges.
## Troubleshooting
| Symptom | Likely cause |
|---|---|
| Every request denied | Ratelimit service down/unready, or `cpln/rateLimitAddress` wrong — enforcement is fail-closed |
| Requests never limited | `cpln/rateLimitDescriptors` unset (default matches nothing); `domain` mismatch between tag and config; rules edited without `force-redeployment`; misspelled tag (tags are unvalidated) |
| Cloudflare 526 / TLS errors | Domain TLS not using the origin-cert secret, or SSL mode not Full (strict) |
| Redirect loop behind Cloudflare | SSL mode is Flexible — switch to Full (strict) |
| Origin still reachable directly | `inboundAllowCIDR` not restricted to the CDN ranges |
## Quick reference
- `update_workload` — `cpln/rateLimit*` tags; `inboundAllowCIDR` lock-down
- `create_domain` / `set_domain_tls` / `add_domain_route` — CDN domain wiring (the TLS and ratelimit secrets are managed by the user)
- `list_deployments` — canonical endpoint + readiness checks
- CLI only: `cpln apply --file rate-limiting.yaml` (bundled manifest), `cpln workload force-redeployment` (config reload)
## Related skills
| Need | Skill |
|---|---|
| Workload types, defaults, canonical URL rules — start here | `workload` |
| Firewall rules, CIDR allow-lists | `firewall-networking` |
| TLS, probes, hardening | `workload-security` |
## Documentation
- [Configure CDN Guide](https://docs.controlplane.com/guides/configure-cdn.md)
- [Rate Limiting Guide](https://docs.controlplane.com/guides/rate-limiting.md)
- [Secret Reference (TLS)](https://docs.controlplane.com/reference/secret.md)
cpln23.9 KB
---
name: cpln
description: "Writes cpln CLI commands and workflows for Control Plane. Use when the user asks about cpln login, cpln apply, cpln workload, CLI or CI/CD deploys, container debugging with cpln exec/logs, or any cpln resource command."
---
# cpln CLI
**MCP first; the CLI is the fallback** — use it when the MCP server is unavailable or unauthenticated, for the CLI-only operations below, and for interactive debugging or scripted GitOps. **In CI/CD the CLI is the primary interface** — pipelines authenticate with a service-account key in `CPLN_TOKEN`, build and push images (`cpln image build --push`, or `--remote` on a runner with no Docker daemon), and apply resources (`cpln apply --ready`). Platform rules (resource model, secrets, destructive ops, production defaults, scale-to-zero, firewall) live in the operating guide (`get_cpln_rules`); this skill is the CLI mechanics.
**Never write a `cpln` command from memory.** Verify every verb and flag with `cpln <command> --help` before quoting it. If a command isn't in the resource command map below, assume it isn't real.
## CLI-only operations
No MCP equivalent — the CLI's primary job:
| Command | Purpose |
|---|---|
| `cpln image build` | Build a container image — locally with Docker, or on Control Plane with `--remote` (no daemon) — and push to the org registry |
| `cpln image copy` | Copy an image between orgs |
| `cpln port-forward` | Forward local ports to a running workload |
| `cpln convert` | Convert Kubernetes manifests to Control Plane specs (Compose is `cpln stack`) |
| `cpln cp` | Copy files in or out of a running container |
| `cpln apply` | Scripted GitOps — declarative create-or-update from files |
One-shot and interactive container commands, TTY sessions (`workload connect`, `exec -it`), and streamed logs (`cpln logs --tail`) are also CLI-only; MCP's `get_workload_logs` covers bounded log fetches. Everything else — discovery and CRUD — prefer the MCP tools (generic `list_resources` / `get_resource` / `delete_resource` with a `kind`, typed `create_*`/`update_*` for mutations).
When no MCP tool covers a resource, field, or sub-endpoint, use the `cpln` CLI for that piece (ground the command in this skill and `--help`) or tell the user what is missing. The raw-API escape hatch (`cpln_api_request`) is disabled by default — it bypasses the typed tools' pre-call validation.
## Setup & auth
```bash
cpln login # interactive (opens a browser); creates the "default" profile
cpln profile update default --org ORG --gvc GVC # set defaults ("update" creates the profile if missing)
```
**CI/CD needs no profile.** With `CPLN_TOKEN` set (service-account key, from `cpln serviceaccount add-key`), the CLI runs a profile-less session against `api.cpln.io`. Add `CPLN_ORG` / `CPLN_GVC` for defaults and `CPLN_SKIP_UPDATE_CHECK=1` to silence update checks. Resolution everywhere is **flag, then env var, then profile**: `--org` beats `CPLN_ORG` beats the profile default — same for `--gvc`/`CPLN_GVC`, `--profile`/`CPLN_PROFILE`, `--endpoint`/`CPLN_ENDPOINT`, `--token`/`CPLN_TOKEN`. Profiles live in `~/.config/cpln` (override with `CPLN_HOME`).
**Never pass `--token`** — it leaks into logs and shell history; use `CPLN_TOKEN` or a profile. Inspect context with `cpln profile get` — **there is no `cpln whoami`.** **`cpln profile token` (prints the profile's live access JWT) is break-glass** — it exposes a live credential: never suggest it or run it on your own; use it only when the user explicitly asks. **Secret data commands are off-limits entirely** — never run or suggest `cpln secret reveal`, `cpln secret create-*`, `cpln secret edit`, or `cpln secret delete`; the user manages secret values and lifecycle themselves. Explain any profile state changes so operators can revert them.
## Command structure & shared flags
```
cpln <resource> <action> [REF] [--flags]
```
Standalone (break the pattern): `cpln apply`, `delete`, `logs`, `port-forward`, `cp`, `convert`, `login`. Aliases: `workload`=`w`, `identity`=`id`, `serviceaccount`=`sa`, `location`=`loc`, `stack`=`compose`. Shell completion: `cpln misc install-completion` (bash/zsh/fish).
Flags on nearly every command — never list per-command:
- **Context**: `--profile`, `--org`, `--gvc`
- **Output**: `--output`/`-o` (`text|json|yaml|json-slim|yaml-slim|tf|crd|names`), `--color`, `--ts` (`iso|local|age`), `--max` (default 50; `0` = all)
- **Request**: `--token`, `--endpoint`, `--insecure`/`-k` · **Debug**: `--verbose`/`-v`, `--debug`/`-d`
- **Always use `yaml-slim`/`json-slim` for round-tripping.** Plain `yaml`/`json` include server-side fields (`status`, `id`, `created`, `lastModified`, `links`) that break `cpln apply`.
- **Include `--org` (and `--gvc`) explicitly on every mutation**, even with profile defaults. `--gvc` exists on all subcommands of GVC-scoped resources (workload, identity, volumeset) plus helm, stack, apply, convert, cp, delete, port-forward — but **not** on `cpln logs` (GVC goes inside the LogQL query).
- **Cap lists with `--max`.** Omit only when targeting a specific named resource.
## Standard CRUD
| Action | Syntax | Notes |
|---|---|---|
| **List** | `cpln <resource> get` | No args = list all. **There is NO `list` subcommand.** |
| **Get** | `cpln <resource> get REF...` | |
| **Create** | `cpln <resource> create --name NAME` | Also `--description`, `--tag K=V` |
| **Delete** | `cpln <resource> delete REF...` | Multiple refs |
| **Edit** | `cpln <resource> edit REF` | Opens YAML in `$EDITOR`. `--replace` replaces instead of merging |
| **Patch** | `cpln <resource> patch REF --file FILE` | |
| **Tag** | `cpln <resource> tag REF... --tag K=V` | Remove: `--remove-tag KEY` |
| **Update** | `cpln <resource> update REF --set PROP=VAL` | Also `--unset PROP`; array props take `+=` / `=` / `-=` |
| **Clone** | `cpln <resource> clone REF --name NEW` | Spec only. Not on every kind (map below) |
| **Audit** | `cpln <resource> audit [REF]` | `--since` (default 7d), `--from`/`--to`, `--subject`, `--context` |
| **Query** | `cpln <resource> query` | See below |
Also on most kinds: `access-report REF`, `eventlog REF`, `permissions` (no args). **`eventlog` has the alias `log`** — `cpln workload log` shows platform events, NOT container logs (those come from `cpln logs`).
### Query
```bash
cpln workload query --match all --tag environment=production --tag region=europe
cpln workload query --match any --rel gvc=gvc-a --rel gvc=gvc-b
cpln workload query --property name=my-workload
```
`--match` (`all` default / `any` / `none`); `--tag KEY=VALUE`, `--property`/`--prop NAME=VALUE` (e.g. `status.phase=running`), `--rel KIND=VALUE` — all repeatable. The `gvc`, `policy`, and `group` create commands accept `--query-match`/`--query-tag`/`--query-property`/`--query-rel` (group also `--query-kind user`) for dynamic targeting. Full language: `query-spec` skill.
## Resource command map
Core anti-hallucination reference. **Scope**: org = needs `--org`; gvc = needs `--org` + `--gvc`; local = no API call. "Full" = create, get, delete, edit, patch, tag, update (audit, eventlog, query, access-report, permissions exist nearly everywhere).
| Resource | Scope | CRUD | Non-standard subcommands |
|---|---|---|---|
| **workload** (`w`) | gvc | Full + clone | `connect`, `exec`, `run`, `cron` (get/run/start/stop), `replica` (get/stop), `force-redeployment`, `get-deployments`, `open`, `start`, `stop` |
| **gvc** | org | Full + clone | `add-location`, `remove-location` (both `--location`, repeatable), `delete-all-workloads` |
| **secret** | org | get only — secret data and lifecycle are managed by the user | — |
| **policy** | org | Full + clone | `add-binding`, `remove-binding` |
| **identity** (`id`) | gvc | Full, **no clone** | — |
| **volumeset** | gvc | Full, no clone | `expand`, `shrink`, `snapshot` (create/delete/get/restore), `volume` (delete/get) |
| **domain** | org | create, delete, edit, patch, tag — **no update**, no clone | — |
| **cloudaccount** | org | **No generic create, no update** | `create-aws`, `create-azure`, `create-gcp`, `create-ngs` |
| **image** | org | get, delete, edit, patch, tag only | `build`, `copy`, `docker-login` |
| **agent** | org | Full, no clone | `info`, `manifest`, `up` |
| **group** | org | Full + clone | `add-member`, `remove-member` (`--email`, `--serviceaccount`) |
| **ipset** | org | Full + clone | `add-location`, `remove-location`, `update-location` |
| **serviceaccount** (`sa`) | org | **No update**; clone | `add-key` (`--description` required), `remove-key` |
| **mk8s** | org | **No create**; clone | `dashboard`, `health`, `join`, `kubeconfig` |
| **user** | org | No create | `invite` (`--email`, optional `--group`) |
| **org** | — | create (needs `--accountId`, `--invitee`); **no delete** | — |
| **profile** | local | get, delete, update (creates if missing; alias `create`) | `login`, `set-default`, `token` |
| **helm** | gvc | — | `install` (alias `apply`), `upgrade`, `uninstall`, `get`, `list`, `history`, `rollback`, `template` |
| **stack** (`compose`) | gvc | — | `deploy` (alias `up`), `manifest`, `rm` (alias `down`) |
| **location** (`loc`) | org | create, delete (BYOK locations only), edit, patch — no update | `install`, `uninstall` |
| **auditctx** | org | Full + clone, **no delete** | — |
| **quota** | org | get, edit, patch | — |
| **task** | org | get, delete | `complete`, `get-mine` |
| **account** | — | get only | — |
| **rest** | — | — | `get`, `post`, `put`, `patch`, `delete`, `create`, `edit` against raw API paths |
| **operator** | local | — | `install`, `uninstall` |
## Non-standard commands
### cpln apply / cpln delete
```bash
cpln apply --file ./manifests/ --gvc GVC --ready # DIRECTORY — recursive over .yaml/.yml/.json
cpln apply --file all.yaml --ready # MULTI-DOC file (resources split by ---)
cpln apply --file - < manifest.yaml # stdin
cpln delete --file manifest.yaml # delete the resources listed in a file
```
**Apply multi-resource deploys in one call** (directory or multi-doc file) — the CLI sorts resources into dependency order before applying: `agent, secret, cloudaccount, gvc, identity, volumeset, policy, workload` (other kinds after); `cpln delete --file` runs the same order reversed. Splitting into multiple calls reintroduces the ordering problem. Apply is a PUT upsert — it prints `Created`/`Updated` per resource. `--ready` polls workloads (5s interval, up to 5 min) until ready. `--k8s` converts Kubernetes manifests inline (Deployment/Secret/ConfigMap/PVC; pull secrets get linked onto the GVC).
GVC targeting for GVC-scoped resources:
- **Single-GVC bundle** (common): pass `--gvc GVC` — it fills in the GVC for every resource that doesn't declare one.
- **Multi-GVC bundle**: declare `gvc:` inline as a top-level field (same level as `kind`/`name`) and omit the flag.
- Inline `gvc:` plus a **different** `--gvc` value = hard error; they must agree.
- Org-scoped resources ignore the GVC field/flag entirely.
```yaml
kind: workload
name: my-app
gvc: prod # target GVC declared inline
spec: { ... }
```
### cpln logs
```bash
cpln logs '{gvc="GVC", workload="WORKLOAD"}' --org ORG --tail
```
- The query is a **positional argument** (first arg, single quotes, LogQL). `--gvc` is **not** a flag here.
- Labels: `container`, `gvc`, `location`, `provider`, `replica`, `stream`, `workload`. Special: `container="_accesslog"` for HTTP access logs.
- Filters: `|= "error"` (contains), `!= "debug"` (excludes), `|~ "timeout|crash"` (regex).
- Streaming: `--tail` (also `-t`/`-f`). **`--follow` does NOT exist.**
- `--limit N` (default 30; `0` = unlimited, auto-paginates), `--since` (default `1h`), `--from`/`--to`, `--direction forward|backward`, and its own `-o default|raw|jsonl` (`raw` strips labels and timestamps). Full LogQL: `logql-observability` skill.
### cpln workload create
```bash
cpln workload create --name APP --image IMAGE --gvc GVC [flags]
```
`--type` (`serverless|standard`, default `standard`) — **`stateful` and `cron` CANNOT be created via CLI flags; use `cpln apply --file`.** Other flags: `--port` (default 8080 — must match the container's listening port), `--public`, `--identity`, `--env KEY=VALUE`, `--cpu` (default 50m), `--memory`/`--mem` (default 128Mi), `--volume`, `--container-name`, `--inherit-env`. Internal images: `//image/NAME:TAG`.
### Debugging — exec / connect / run / cron / replica
- **exec** — one-shot command in an existing replica: `cpln workload exec APP --gvc GVC -- ls -la`. **The `-- CMD ARG1 ARG2...` part must be last on the line** — every cpln flag (`--container`, `--location`, `--replica`, `--stdin`/`-i`, `--tty`/`-t`, `--quiet`/`-q`) goes before the `--`; everything after it runs in the replica (same rule for `run` and `cron run`)
- **connect** — interactive shell: `cpln workload connect APP --gvc GVC` (`--shell`, default `bash`; same targeting flags)
- **run** — temporary workload + command: `cpln workload run --image IMAGE --gvc GVC -- CMD` (`--clone WORKLOAD`, `--rm`, `-i`, `--cpu`, `--memory`, `--command`/`-c`, `--arg`/`-a`, `--location`)
- **cron run** — one-off execution of a cron workload: `cpln workload cron run --gvc GVC -- CMD` (`--background`/`-b`, `--timeout` default 600s, `--identity`, `--image`, `--env`)
- **cron start** — trigger the job now, optionally overriding `--env`, `--command`, `--arg`, `--active-deadline-seconds`; **cron stop** REF needs `--replica-name` + `--location` (both required); **cron get** REF lists job executions
- **replica get / stop** — list replica names per location; `stop` requires `--replica-name` + `--location`
When `--location` / `--replica` / `--container` are omitted, the CLI defaults to the GVC's first location, the first replica, and the only container (multi-container: the first one with `ports`, else the first) — it prints a "defaulting to" notice. Pass all three explicitly on multi-location or multi-container workloads.
### cpln policy add-binding
```bash
cpln policy add-binding POLICY --permission reveal --identity //gvc/GVC/identity/ID
```
`--permission` required; at least one principal flag required; **all repeatable** (`--email`, `--serviceaccount`, `--group`, `--identity` — name or full link).
### cpln port-forward / cp
```bash
cpln port-forward WORKLOAD [LOCAL:]REMOTE... --gvc GVC # --address (default localhost), --location, --replica
cpln cp LOCAL WORKLOAD:PATH --gvc GVC # reverse the args (WORKLOAD:PATH LOCAL) to copy out; --container, --location, --replica
```
### Migration tools
- `cpln convert --file K8S.yaml` — Kubernetes manifest to Control Plane spec (`--protocol http|http2|grpc|tcp` for container ports)
- `cpln helm install RELEASE CHART --gvc GVC` — Helm charts and template catalog installs (`--wait`, `--timeout` default 300s, `--set`, `--values`)
- `cpln stack deploy --gvc GVC` — Docker Compose from the current directory (`--dir`, `--compose-file` for alternative naming, `--build` default true)
### Volumeset command verbs
Dedicated verb per operation; the flags are **singular** — `--location`, `--volume-index` (plural forms don't exist):
| Operation | Command | Risk |
|---|---|---|
| Expand volume | `cpln volumeset expand REF --new-size GIB [--location LOC] [--volume-index N]` | Safe |
| Create snapshot | `cpln volumeset snapshot create REF --snapshot-name NAME [--location LOC] [--volume-index N]` | Safe |
| Restore snapshot | `cpln volumeset snapshot restore REF --snapshot-name NAME --location LOC --volume-index N` (all required) | Overwrites volume state |
| Delete snapshot | `cpln volumeset snapshot delete REF --snapshot-name NAME` | Destructive |
| Delete volume | `cpln volumeset volume delete REF [--location LOC] [--volume-index N]` | Destructive (data loss) |
| Shrink volume | `cpln volumeset shrink REF --new-size GIB [--location LOC]` | **DESTRUCTIVE — permanent data loss** |
`shrink` provisions a new, smaller volume and removes the old one — data is **not** migrated. Safe only with built-in redundancy (Kafka replication; Cassandra/CockroachDB), on `ext4`/`xfs` (not `shared`). Apply the destructive-op confirmation from the operating guide (`get_cpln_rules`) first. Detail: `stateful-storage` skill.
## Commands that don't exist
| Wrong | Correct |
|---|---|
| `cpln <resource> list` | `cpln <resource> get` (no args = list all) |
| `cpln logs --follow` | `cpln logs --tail` (or `-t` / `-f`) |
| `cpln workload log` for container logs | that's the `eventlog` alias (platform events); container logs = `cpln logs '{gvc="GVC", workload="W"}'` |
| `cpln cloudaccount create` | `cpln cloudaccount create-aws` / `create-azure` / `create-gcp` / `create-ngs` |
| `cpln mk8s create` | `cpln apply --file mk8s-manifest.yaml` |
| `cpln workload update REF --identity X` | `cpln workload update REF --set spec.identityLink=//identity/X` |
| `cpln gvc update REF --location LOC` | `cpln gvc add-location REF --location LOC` |
| `cpln identity clone` | identity has no clone — `get -o yaml-slim`, edit the name, `cpln apply` |
| `cpln volumeset ... --locations / --volume-indexes` | singular: `--location`, `--volume-index` |
| `cpln whoami` | `cpln profile get` (context) / `cpln account get` |
## Building & referencing images
`cpln image build` puts an image in the org's **private registry** — the registry referenced in workload specs as `//image/NAME:TAG`. Two ways to run it:
```bash
cpln image build --name my-app:v1.0 --remote # builds on Control Plane, pushes for you — no Docker
cpln image build --name my-app:v1.0 --push # builds through the local Docker daemon
```
- **`--remote`**: uploads `--dir` (default `.`, filtered by `.dockerignore` or `.gitignore`, capped at 500 MB / 20,000 files), or builds a GitHub/GitLab HTTPS repo given with `--repo` (+ `--branch`). The service picks Dockerfile vs generated build and always produces `linux/amd64`. `--detach` returns a build id; `Ctrl+C` stops watching but not the build.
- **Local**: Dockerfile auto-detected in `--dir` or passed with `--dockerfile PATH`, built with Docker (buildx when available); no Dockerfile means a buildpack build via `pack` (auto-downloaded), default builder `heroku/builder:24_linux-amd64`, everything after `--` goes to `pack`. `--platform` defaults to `linux/amd64`; multi-arch lists need buildx and `--push`. `--env` is **build-time only**, never runtime.
- **`--dockerfile`, `--builder`, `--buildpack`, `--env`, `--env-file`, `--trust-builder`, `--trust-extra-buildpacks`, `--platform`, and `--push` are rejected with `--remote`**; `--repo`, `--branch`, and `--detach` error without it.
- **A Docker daemon is required for a local build only.** On thin CI runners use `--remote`, or `docker build` + `docker push` against the org registry after `cpln image docker-login`.
- Cross-org copy: `cpln image copy NAME:TAG --to-org ORG2 [--to-name NEW] [--to-profile P]` — logs into both registries, then pulls, tags, pushes. **`copy` has no `--remote` mode; it still needs a local daemon.**
Image reference rules in workload specs:
| Source | Reference in spec | Pull secret? |
|---|---|---|
| **Your org's registry (internal)** | `//image/NAME:TAG` — never the `ORG.registry.cpln.io` hostname | No |
| **Public Docker Hub** | bare name (`nginx:latest`) — never add `docker.io/` | No |
| **Other public registry** | exact host path (`gcr.io/...`, `ghcr.io/...`) | No |
| **External private registry** (ECR, GCR, ACR, private Docker Hub, another CPLN org) | exact host path | **Yes** — `docker`/`ecr`/`gcp` secret on GVC `spec.pullSecretLinks` |
Full image workflow (remote vs local build detail, buildx fallback, cross-org copy, pull-secret setup): `image` skill.
## Workflow: Deploy a workload
The CLI runbook (also the CI/CD path):
```bash
cpln gvc create --name my-gvc --location aws-us-west-2 --org my-org
cpln image build --name my-app:v1.0 --push # or --remote, with no Docker daemon
cpln workload create --name my-app --gvc my-gvc --image //image/my-app:v1.0 --port 8080 --public
cpln workload get-deployments my-app --gvc my-gvc # verify readiness
```
## Workflow: Grant secret access (3 steps)
The 3-step rule (identity + policy + reference) is owned by the operating guide (`get_cpln_rules`). The secret must already exist (created by the user — offer to draft the manifest for them to fill and apply; `setup-secret` skill). CLI fallback:
```bash
cpln identity create --name my-app-identity --gvc my-gvc --org my-org
cpln workload update my-app --gvc my-gvc --set spec.identityLink=//identity/my-app-identity
cpln policy create --name secret-access --target-kind secret --resource db-password --org my-org
cpln policy add-binding secret-access --permission reveal \
--identity //gvc/my-gvc/identity/my-app-identity --org my-org
cpln workload update my-app --gvc my-gvc \
--set spec.containers.main.env.DB_PASSWORD.value=cpln://secret/db-password.payload
```
## Workflow: GitOps with cpln apply
```bash
cpln gvc get my-gvc -o yaml-slim > manifests/gvc.yaml # export (slim strips server fields)
cpln workload get my-app --gvc my-gvc -o yaml-slim > manifests/workload.yaml
cpln apply --file ./manifests/ --gvc my-gvc --ready # idempotent; run on every push
```
## Workflow: Rename or clone (names are immutable)
No `rename` exists. **Preferred — `clone`** (on workload, gvc, policy, group, ipset, serviceaccount, mk8s, auditctx; duplicates spec only; secrets are off-limits — the user clones them):
```bash
cpln workload clone old-name --name new-name --gvc my-gvc
cpln workload get-deployments new-name --gvc my-gvc # verify healthy
cpln workload delete old-name --gvc my-gvc # only after verification
```
Kinds without clone (identity, volumeset, domain, agent): `get -o yaml-slim`, edit the name, `cpln apply`. Renaming a workload changes its internal hostname (`WORKLOAD.GVC.cpln.local`) and public URL — update domain routes (`spec.ports[].routes[].workloadLink`), policy `targetLinks`/`targetQuery`, internal-DNS callers, and external clients. Never delete the old workload before the new one is verified healthy.
## Workflow: Debug a failing workload
```bash
cpln logs '{gvc="my-gvc", workload="my-app"}' --org my-org --tail # stream logs
cpln logs '{gvc="my-gvc", workload="my-app"} |= "error"' --org my-org # filter (LogQL, not a shell pipe)
cpln workload exec my-app --gvc my-gvc -- ls /app # inspect the container filesystem
cpln workload connect my-app --gvc my-gvc # interactive shell
cpln port-forward my-app 8080:8080 --gvc my-gvc # probe locally
```
**Cron workloads — query logs per execution, not per workload.** A plain `{gvc=, workload=}` query mixes every past run. Enumerate executions with `cpln workload cron get NAME --gvc GVC`, then scope the LogQL with the `replica` label plus a time window. Full pattern: `logql-observability` skill.
## Platform rules & integration
Scale-to-zero/autoscaling, production defaults/probes, Template Catalog first, destructive ops, secrets, and firewall rules live in the operating guide (`get_cpln_rules`) and their dedicated skills — not duplicated here. Before authoring any apply YAML / CI manifest / API body, call `get_resource_schema`. IaC: Terraform (`controlplane-com/cpln` provider), Pulumi (`@pulumiverse/cpln`), K8s Operator (`cpln operator install`).
## Related skills
| Need | Skill |
|---|---|
| Remote vs local build detail, buildx fallback, pull secrets | `image` |
| LogQL beyond the basics, per-execution cron queries | `logql-observability` |
| Query language (`--match` / `--tag` / `--rel`) | `query-spec` |
| Volumeset semantics and shrink safety | `stateful-storage` |
| K8s / Compose / Helm migration | `migration-patterns` |
| Pipelines and GitOps patterns | `gitops-cicd` |
| Terraform / Pulumi / K8s operator | `iac-terraform-pulumi`, `k8s-operator` |
## Documentation
- Platform guardrails & resource model: the operating guide (`get_cpln_rules`)
- [Control Plane Docs](https://docs.controlplane.com) · [AI page index](https://docs.controlplane.com/llms.txt) · [CLI Reference](https://docs.controlplane.com/cli-reference/overview.md)
domain12 KB
---
name: domain
description: "Custom domains for Control Plane workloads. Use when the user asks to put a domain or subdomain in front of a workload, pick cname vs ns, configure routing or TLS, or hits apex, ownership, or workloadLink errors."
---
# Custom Domains
A `domain` is an org-level resource that binds a DNS name to workloads in **one GVC**. **Created ≠ live:** after the resource exists, the user still adds records at their DNS provider — read exactly which from `status.dnsConfig` and hand them over verbatim, never guessed. Every shape decision below is platform-enforced and a wrong combination is a rejected mutation, so decide BEFORE calling `create_domain` (the tool requires `dnsMode` and `ports` explicitly). Never set `spec.domain` on a GVC — that legacy field is deprecated; the Domain resource is the only path.
## Decide the shape first
**1. Apex or subdomain?** The apex is the registrable root (`example.com`, `example.co.uk`); anything deeper is a subdomain (`app.example.com`).
**2. `dnsMode` — who runs DNS:**
| Mode | Valid for | Wiring | Cert challenge |
|---|---|---|---|
| `cname` | **apex (required)** and subdomains | User adds CNAME records per `status.dnsConfig` | `http01` default, `dns01` opt-in |
| `ns` | **subdomains only** | Delegates the subdomain zone via 4 NS records (`ns1`/`ns2.cpln.cloud`, `ns1`/`ns2.cpln.live`) | `dns01` only — `http01` rejected |
`dnsMode` defaults to `cname` (to `ns` when `gvcLink` is set). The platform rejects `ns` on an apex, and rejects a `cname` domain nested under an existing NS domain (`parent_ns_domain_exists`).
**3. Routing — exactly ONE of three.** All routes in a domain must target workloads in the **same GVC**.
| Mode | What it does | Constraints |
|---|---|---|
| `ports[].routes` | Explicit path routes to workloads | The default choice; works for every workload type |
| `gvcLink` | Every workload in the GVC gets `{workload}.{domain}` | Excludes `workloadLink` and any `ports[].routes`. With `cname` + `http01` it demands `tls.serverCertificate` on every TLS port (http01 cannot issue wildcard certs) |
| `workloadLink` (spec-level) | Replica-direct: binds the whole domain to ONE stateful workload with per-replica DNS names | **Stateful only** (`workloadLink must link to a stateful workload`); every port exactly ONE route to that same workload; `http01` rejected |
For an app, site, or API on serverless/standard, the answer is `ports[].routes`. Route-level `workloadLink` inside `routes[]` is a different field with no stateful restriction.
## Ownership and create order
- **`ns` subdomain:** the apex domain resource must already exist in the org (`apex_must_exist`).
- **Everything else:** ownership is proven either by the org already owning the verified apex (subdomains then attach with no extra records), or by a TXT record — the create fails with `must_prove_ownership` listing the options: `_cpln.{apex}` / `_verify.{apex}`, or `_cpln-{label}.{rest}` / `_verify-{label}.{rest}` at any segment level, value = org GUID **or** org name (TTL 600). The user adds **one**, waits for propagation, and you retry the same create. `create_domain` surfaces these records in its error output.
- **Apex owned by another org?** The apex name itself is taken (globally unique), but **subdomains still work**: they go through the same TXT proof in this org — the standard multi-org pattern (keep the apex in the production org).
- **`.internal` domains** are strict same-org — apex and subdomains must live in one org (`apex_owned_by_other_org`, HTTP 409) — and: `cname` only, no `gvcLink`, `certChallengeType` forbidden; every TLS port needs `tls.serverCertificate.secretLink` (no ACME).
## Manifest shape
```yaml
kind: domain
name: app.example.com
spec:
dnsMode: cname
ports: # max 10 per domain
- number: 443 # default 443; 443 + http/http2 auto-gets a TLS block
protocol: http2 # http | http2 | tcp (tcp needs a dedicated load balancer)
routes: # max 150 per port (200 with tag cpln/routeLimitOverride)
- prefix: /api # prefix XOR regex (RE2); prefix defaults to "/"
replacePrefix: / # optional rewrite before forwarding
workloadLink: //gvc/GVC/workload/API
port: 8080 # optional target container port
- prefix: /
workloadLink: //gvc/GVC/workload/FRONTEND
```
- **Longest prefix wins** — prefix routes are auto-sorted; regex routes are NOT sorted, written order matters. Duplicate prefix+host combinations are rejected (`There are more than one routes for the prefix …`).
- **Listener ports other than 443/80 — and the `tcp` protocol — require a dedicated load balancer.** Without one the domain deploys into `warning` (`Unable to configure port …`) instead of serving.
- **Subdomain matching on one domain** (`hostPrefix` / `hostRegex`, mutually exclusive) requires `acceptAllHosts` or `acceptAllSubdomains` (which exclude each other) AND a GVC with a dedicated load balancer. `hostPrefix` charset: alphanumeric, dot, underscore, hyphen.
- **Header rewrites** (`headers.request.set`): values may use only `%REQUESTED_SERVER_NAME%`, `%DOWNSTREAM_REMOTE_ADDRESS_WITHOUT_PORT%`, `%START_TIME%`.
- **Traffic mirroring** per route: `mirror: [{workloadLink, percent 0-100, port}]` — same GVC, response comes only from the primary.
- **CORS** per port: `allowOrigins` entries take `exact` XOR `regex`; header lists are lowercased; `maxAge` format is digits + `h`/`m`/`s` only.
- **TLS** per port: `minProtocolVersion` default `TLSV1_2`; custom cert = keypair secret (PEM) on `serverCertificate.secretLink`; `clientCertificate` enables mTLS verification — client cert details reach the workload in the XFCC header.
## Certificates
- Let's Encrypt, auto-provisioned for port 443 once validation passes; ~90-day certs renewed automatically.
- `cname` defaults to `http01`: DNS must already resolve and `/.well-known/acme-challenge/` must redirect to the platform solver (`http01-solver.cpln.io`) — a CDN/WAF forcing HTTPS or blocking the path breaks it; switch to `dns01`. `ns` always uses `dns01`. The tag `cpln/skipDNSCheck: "true"` skips the DNS-propagation gate in certificate processing.
- **`dns01` on a `cname` domain adds an extra record**: a `_acme-challenge.{host}` CNAME appears in `status.dnsConfig` — without it the certificate never issues.
- Wildcard certs come only from `dns01` — that is why `cname` + `gvcLink` + `http01` demands a custom certificate.
## After create — DNS records and status
- Read the domain back and give the user the records from `status.dnsConfig`. CNAME mode points at the GVC endpoint alias. Via the MCP tools the alias is resolved for you, so the CNAME target comes back ready to paste (e.g. `0p2fpmbe7sr5c.t.cpln.app`); via the `cpln` CLI the value is the literal `<gvcAlias>.t.cpln.app` placeholder — substitute the GVC's top-level `alias` field. **Never hand the user a `<gvcAlias>` placeholder as a DNS record.** `cname` + `gvcLink` needs one CNAME per workload; `workloadLink` adds per-replica records (`{workload}-{i}-{location}`).
- Many DNS providers refuse CNAME at the apex — the user needs ALIAS/ANAME support or a CDN in front.
- `status.status`: `initializing`, `pendingDnsConfig` (records not seen yet), `pendingCertificate` (validated, cert issuing), `ready`; `warning`/`errored` carry detail in `status.warning`; `usedByGvc` marks a domain referenced by the legacy GVC `spec.domain`. Pending states are expected, not errors. Misconfigurations that pass schema validation — routes to a missing GVC/workload, no valid routes, a disallowed port/protocol, an ignored `hostPrefix` — land as `warning` and increment the `domain_warnings` metric.
- **Host header:** serverless workloads receive the canonical endpoint as `Host` (the custom domain arrives in `X-Forwarded-Host`); standard/stateful receive the custom domain.
- Report honestly: "domain created, routes configured, DNS records pending at your provider" + the record list. Never claim the domain is serving before DNS exists.
## Platform rejections and exact fixes
| Error | Fix |
|---|---|
| `cname is the only valid dnsMode for apex domain X` | Use `dnsMode: cname` on the apex; `ns` only delegates a subdomain zone |
| `The apex domain X must be created before Y` | Create the apex domain resource first, then the subdomain |
| `apex_owned_by_other_org` (409) | `.internal` only — internal apex + subdomains stay in one org. A public subdomain under another org's apex just needs the TXT proof |
| `must_prove_ownership` | Hand the user ONE TXT record from the response, wait for propagation, retry |
| `parent_ns_domain_exists` | A CNAME domain cannot live under an NS domain — create it as part of the NS zone or restructure |
| `workloadLink must link to a stateful workload` | Drop spec-level `workloadLink`; route serverless/standard via `ports[].routes` |
| `Only one of gvcLink or ports.routes may be configured` | Pick one routing mode |
| `when workloadLink is configured, every port must have exactly ONE route` / `no route can reference another workload` | One route per port, all to the linked workload — or drop `workloadLink` |
| `certChallengeType can not be http01` / `http01 … not supported for dnsMode ns` | Use `dns01` or omit `certChallengeType` |
| `Domains may only route to Workloads in a single GVC` | Split into one domain per GVC, or move the workloads |
| `hostPrefix or hostRegex can only be used if …` | Set `acceptAllHosts` or `acceptAllSubdomains` (and use a dedicated load balancer) |
| `number of routes exceeds maximum of 150` | Consolidate routes, or add tag `cpln/routeLimitOverride` (raises to 200) |
## Verify
1. `get_resource` (kind `domain`) — `status.status` progressing, `status.dnsConfig` matches what the user added.
2. After the user adds records: `dig TXT _cpln.DOMAIN` / `dig CNAME DOMAIN` to confirm propagation before retrying or polling.
3. Once `ready`: `curl -I https://DOMAIN/PATH` and confirm each prefix lands on the intended workload.
## Quick reference — MCP tools
| Tool | Action |
|---|---|
| `create_domain` | Create — `dnsMode` and `ports` required; pre-validates apex/exclusivity rules; surfaces ownership TXT records on failure |
| `update_domain` | Description/tags, `acceptAll*` flags, `gvcLink`/`workloadLink` bind or remove. CANNOT touch ports, dnsMode, certChallengeType |
| `add_domain_port` / `remove_domain_port` | Add a listener (errors if the number exists) / remove one (destructive — live traffic on that port stops) |
| `add_domain_route` / `update_domain_route` / `remove_domain_route` | Manage routes on a port; update/remove identify the route by `routeIdentifier` (`prefix` or `regex`); removal 404s matched traffic until re-routed |
| `set_domain_tls` / `clear_domain_tls` | Overwrite or remove the whole TLS block on a port — on 443 with http/http2 the default TLS block comes back (TLS cannot be disabled there) |
| `set_domain_cors` / `clear_domain_cors` | Overwrite or remove the whole CORS block on a port |
| `get_resource` / `list_resources` / `delete_resource` (kind `domain`) | Read / list / delete — names are FQDNs, passed as-is; delete is destructive, confirm first |
CLI fallback (read the `cpln` skill first; CI/CD = `CPLN_TOKEN` + `cpln apply`): `cpln domain create` takes only `--name`/`--description`/`--tag` — spec changes go through `cpln domain edit` or `cpln domain get -o yaml-slim` + `cpln apply`. There is no `cpln domain update`.
## Related skills
| Need | Skill |
|---|---|
| Workload ports, exposure, canonical URL | `workload` |
| Dedicated load balancer (wildcard hosts, tcp ports) | `ipset-load-balancing` |
| CDN/WAF in front, rate limiting | `cdn-rate-limiting` |
| Keypair secrets for custom certificates | `setup-secret` |
## Documentation
- [Domain Reference](https://docs.controlplane.com/reference/domain.md)
- [Configure a Domain Guide](https://docs.controlplane.com/guides/configure-domain.md)
- [Custom Domain Quickstart](https://docs.controlplane.com/quickstart/quick-start-3-custom-domain.md)
- [cpln domain CLI](https://docs.controlplane.com/cli-reference/commands/domain.md)
environment-promotion9.41 KB
---
name: environment-promotion
description: "Promotes workloads across dev/staging/production on Control Plane. Use when the user asks about environment promotion, org-per-environment, cross-org image pulls, image promotion, deploying to production, or rollback."
---
# Environment Promotion
Control Plane has **no built-in promote or rollback primitive** — promotion is applying the same artifacts (image + manifests) to the next environment. Two topologies exist: **org-per-environment (the documented best practice)** and GVC-per-environment. The recurring failure is image access: a staging/prod org cannot pull the dev org's images until you either copy the image or wire up a cross-org pull secret.
## Choosing a topology
| Topology | Isolation | Image sharing | Best for |
|---|---|---|---|
| **Org per environment** (recommended) | Strongest — policies, secrets, users, audit fully separate | `cpln image copy` or cross-org pull secret | Production, compliance-sensitive teams |
| **GVC per environment** (one org) | Weak — shared org policies and access | Same org registry; no pull secret needed | Small teams, rapid iteration |
With org-per-environment, GVCs and workloads keep **identical names** in every org, so the same manifests apply unchanged — no environment suffixes in resource names. Org creation is account-level: Console, or `cpln org create --accountId ID --invitee EMAIL`.
## Promoting manifests
Keep manifests in git and apply them per environment — `cpln apply` is idempotent (PUT upsert) and resolves resource ordering:
```bash
cpln apply --file ./manifests/ --org my-org-staging --gvc my-gvc --ready # org-per-env: same files, next org
cpln apply --file ./manifests/ --org my-org --gvc staging-gvc --ready # gvc-per-env: same files, next GVC
```
- Bootstrap manifests from a live environment with `cpln <resource> get REF -o yaml-slim` (plain `yaml` output breaks apply).
- Environment differences (env vars, scaling, firewall) belong in the manifests per environment — or patch after apply with `update_workload` / `update_gvc` (both PATCH semantics).
- For IaC-based promotion, export live resources to Terraform: `export_terraform` (one self link, or bulk by path depth — a whole GVC or org), `export_terraform_batch` (full profile, up to 100 explicit links), `convert_to_terraform` (manifest to HCL, dry-run validated). An unsupported kind is rejected with the supported list.
## Sharing images across orgs
### Option A — copy the image (one-time promotions)
```bash
cpln image copy my-app:abc1234 --to-org my-org-prod # same credentials for both orgs
cpln image copy my-app:abc1234 --to-org my-org-prod --to-profile prod-profile --cleanup
```
CLI-only (no MCP tool) and **still requires a running Docker daemon — `copy` has no `--remote` mode** — it docker-logins both registries, then pulls, tags, pushes. `--to-name` renames during copy; `--cleanup` removes the local images (use in CI). Needs `pull` permission on the source image and `create` on images in the target org. After the copy the target references it as `//image/my-app:abc1234` — no pull secret.
### Option B — cross-org pull secret (continuous access)
The target org pulls directly from the source org's registry. Four steps:
1. **Source org — puller credentials**: `add_key_to_service_account` (creates the service account if missing; the key is shown **once**).
2. **Source org — grant pull**: `create_policy` with `targetKind: image`, `targetAll: true` (or `targetQuery` by repository), `addPermissions: ["pull"]`, `addServiceAccounts: [LINK]` — bindings go in the create call.
3. **Target org — docker secret**: have the user create a `docker` secret with this `dockerConfigJson` — offer a manifest scaffold (`data` is this JSON as one string; `setup-secret` skill); the username is the **literal string `<token>`** (the registry rejects anything else; the password is the service-account key):
```json
{ "auths": { "my-org-dev.registry.cpln.io": { "username": "<token>", "password": "SERVICE_ACCOUNT_KEY" } } }
```
4. **Target org — attach to the GVC**: `update_gvc` with `pullSecretLinks: ["//secret/dev-registry-pull"]` (merged with existing), then reference the image by its **full registry hostname** in the workload spec:
```yaml
spec:
containers:
- name: main
image: my-org-dev.registry.cpln.io/my-app:abc1234
```
CLI fallback for the same four steps:
```bash
cpln serviceaccount create --name image-puller --org my-org-dev # CLI does NOT auto-create on add-key
cpln serviceaccount add-key image-puller --description "cross-org pull" --org my-org-dev # save the key
cpln policy create --name image-pull --target-kind image --all --org my-org-dev
cpln policy add-binding image-pull --serviceaccount image-puller --permission pull --org my-org-dev
# the user creates the dev-registry-pull docker secret in my-org-prod, then:
cpln gvc update my-gvc --set 'spec.pullSecretLinks+=//secret/dev-registry-pull' --org my-org-prod
```
Same-org images never need a pull secret — the platform injects a default registry credential for the org's own registry automatically.
## Image tags across environments
- **Promote immutable tags** (git SHA `my-app:abc1234` or semver `my-app:v1.2.3`) — promote the exact artifact you tested; mutable tags (`latest`, `staging`) make rollback unreliable. Digest pins (`my-app@sha256:...`) are maximally reproducible.
- **`supportDynamicTags`** (workload spec, default `false`): redeploys the workload automatically when a tag's underlying digest changes (within ~5 minutes) — useful for dev environments on mutable tags, wrong for production promotion.
## CI/CD promotion pipeline
The CLI is the primary interface in pipelines; `CPLN_TOKEN` alone is enough (no profile needed — see the `cpln` skill). The shape that works:
```yaml
# Build once in dev, then per stage: copy the image + apply the manifests
- run: cpln image build --name my-app:${{ github.sha }} --push # CPLN_TOKEN + CPLN_ORG=my-org-dev
- run: cpln apply --file ./manifests/ --gvc my-gvc --ready
# staging / prod stages (gate each with environment approvals):
- run: cpln image copy my-app:${{ github.sha }} --to-org my-org-prod --cleanup # dev-org token
- run: cpln apply --file ./manifests/ --gvc my-gvc --org my-org-prod --ready # prod-org token
```
- **One service-account token per org** — a dev-org token must not be able to touch prod; the copy step runs with source-org credentials plus a `--to-profile` (or pre-run `cpln image docker-login`) for the target.
- **`--ready` gates promotion** — it polls until workloads are healthy (5s interval, up to 5 min) and fails the job otherwise.
- Approval gates (GitHub environments, GitLab manual jobs) go between stages. Full pipeline setup, npm install (`@controlplane/cli`), runners: `gitops-cicd` skill.
## Rollback
There is no deployment-history rollback — rolling back means **re-pointing the workload at the previous known-good image** (keep the previous tag in git history or your pipeline metadata):
```bash
cpln workload update my-app --set spec.containers.main.image=//image/my-app:v1.1.0 --gvc my-gvc --org my-org
cpln workload get-deployments my-app --gvc my-gvc --org my-org # verify every location reports ready
```
- MCP path: `get_resource` (kind="workload") to record the current image, `update_workload` (`containers: [{name, image}]` — merged by container name), then poll `list_deployments` until ready.
- Org-per-environment: confirm the older image still exists in **this** org's registry first (`list_resources` kind="image") — it may only have been copied forward once.
- **Restart without changing the image**: `cpln workload force-redeployment my-app --gvc GVC` — it PATCHes a `cpln/deployTimestamp` tag with the current time, producing a rolling restart. No MCP equivalent; `update_workload` setting that same tag replicates it.
- Helm-managed releases are the exception with real revision history: `cpln helm rollback RELEASE [REVISION]`.
## Quick reference — MCP tools
| Tool | Purpose |
|---|---|
| `create_gvc` / `create_workload` | Stand up the target environment |
| `update_workload` / `update_gvc` | Patch image, env, scaling, `pullSecretLinks` (PATCH semantics) |
| `add_key_to_service_account` | Puller credentials in the source org (auto-creates the SA; key shown once) |
| `create_policy` | Grant `pull` on images, binding included in the create call |
| `get_resource` (kind `secret`) | Verify the docker pull secret exists before attaching |
| `list_deployments` | Verify a promotion or rollback is ready per location |
| `export_terraform` / `_batch` / `convert_to_terraform` | Export live environments to IaC |
**CLI fallback** (read the `cpln` skill first; CI/CD uses `CPLN_TOKEN` + `cpln apply --ready`): `cpln image copy` is CLI-only and needs a local Docker daemon. `cpln image build --remote` needs none; over MCP only a **repo** build can be started — a local folder must go through the CLI (`image` skill).
## Related skills
| Need | Skill |
|---|---|
| Image building, registries, pull-secret detail | `image` |
| Pipeline setup, runners, service-account auth | `gitops-cicd` |
| Terraform / Pulumi promotion | `iac-terraform-pulumi` |
| Per-environment secrets and RBAC | `access-control` |
## Documentation
- [Environment Promotion Guide](https://docs.controlplane.com/guides/environment-promotion.md)
- [Copy an Image Guide](https://docs.controlplane.com/guides/copy-image.md)
- [cpln apply Guide](https://docs.controlplane.com/guides/cpln-apply.md)
external-logging9.42 KB
---
name: external-logging
description: "Ships Control Plane org logs to external providers. Use when the user asks about log export to S3, CloudWatch, Coralogix, Datadog, Logz.io, Stackdriver, Elastic, syslog, OpenTelemetry, or centralized log forwarding."
---
# External Logging
External logging lives on the **org** (`spec.logging` plus `spec.extraLogging`) and ships **every workload log in the org** — there is no per-GVC or per-workload filtering. One primary provider plus up to 3 extras (4 total); each logging block holds exactly one provider key. Logs stay queryable in built-in LogQL regardless (separate org retention, `spec.observability.logsRetentionDays`, default 30 days). The recurring failure is credentials: each provider needs a pre-created secret of the exact type below — a wrong-type secret passes configuration and the log router then **silently skips that provider**, so logs simply never arrive.
## Providers
| Key | Secret | Required fields | Worth knowing |
|---|---|---|---|
| `s3` | aws | `bucket`, `region`, `credentials` | `prefix` default `/`; region free-form; the IAM user needs `s3:PutObject` on the bucket |
| `cloudWatch` | aws | `region`, `credentials`, `groupName`, `streamName` | `region` is an 18-region allowlist (below); optional `retentionDays` enum and `extractFields` map |
| `coralogix` | opaque | `cluster`, `credentials` | `cluster`: `coralogix.com`, `coralogix.us`, `app.coralogix.in`, `app.eu2.coralogix.com`, or `app.coralogixsg.com`; optional `app`/`subsystem` |
| `datadog` | opaque | `host`, `credentials` | `host` enum: `http-intake.logs.datadoghq.com`, `http-intake.logs.us3.datadoghq.com`, `http-intake.logs.us5.datadoghq.com`, `http-intake.logs.datadoghq.eu` (dashboard `us3.datadoghq.com` pairs with the `us3` intake host) |
| `logzio` | opaque | `listenerHost`, `credentials` | `listenerHost`: `listener.logz.io` or `listener-nl.logz.io` |
| `stackdriver` | gcp | `location`, `credentials` | `location` is a 40-region GCP allowlist (a rejection lists it); the service account needs Logging write |
| `elastic` | aws or userpass | one variant block — see below | |
| `fluentd` | none | `host` | `port` default 24224; Fluent Bit forward protocol |
| `syslog` | none | `host`, `port`, `mode`, `format`, `severity` | `mode` tcp/udp/tls (tls enables TLS); `format` rfc3164/rfc5424; `severity` 0-7; the receiver sees gvc as hostname, workload as appname, replica as procid |
| `opentelemetry` | opaque (optional) | `endpoint` | OTLP over HTTP, not gRPC; an https endpoint turns TLS on; optional `credentials` becomes the Authorization header. Raw header values are intentionally not accepted by the MCP tool. |
- CloudWatch `region` allowlist: us-east-1, us-east-2, us-west-1, us-west-2, ap-south-1, ap-northeast-1, ap-northeast-2, ap-southeast-1, ap-southeast-2, eu-central-1, eu-west-1, eu-west-2, eu-west-3, eu-south-1, eu-north-1, me-south-1, sa-east-1, af-south-1.
- CloudWatch `retentionDays`: 1, 3, 5, 7, 14, 30, 60, 90, 120, 150, 180, 365, 400, 545, 731, 1827, 3653.
### Elastic variants
In YAML the variant is a nested object under `elastic`; the MCP tool flattens it into `elasticVariant` + `indexType` (`indexType` maps to the schema field `type`):
| Variant | Required | Secret |
|---|---|---|
| `aws` | `host` (must end with `es.amazonaws.com`), `port`, `region`, `index`, `type`, `credentials` | aws |
| `elasticCloud` | `cloudId`, `index`, `type`, `credentials` | userpass |
| `generic` | `host`, `index`, `type`, `credentials`; optional `port` (default 443), `path` (must start with `/`) | userpass |
## Configure (MCP first)
1. **Ensure the credential secret exists — but do not pull its value into the chat.** The provider key is the user's own confidential credential: never ask them to paste it here, never pass it as a tool argument, and never invent a placeholder value. Offer to draft the secret manifest with a placeholder for the user to fill and apply (type per the table — usually opaque, `payload` = the raw API key, `encoding: plain`; shapes in `setup-secret`), then **confirm it exists with `get_resource` (kind `secret`) before wiring anything** — referencing a secret that does not exist makes the log router silently skip the provider. No workload identity or policy is needed — the 3-step secret flow applies to workloads consuming secrets, not to org logging.
2. `get_external_logging` — see what is already configured and where.
3. `configure_external_logging`, once per provider. Placement is automatic: with no primary it becomes `spec.logging`; additional providers append to `spec.extraLogging`; re-configuring a provider that is already present updates it in place; a 4th extra errors with "maximum 3 extra logging providers reached". `credentials` takes a bare secret name or `//secret/NAME`. Syslog `mode`/`format`/`severity` are optional here — the tool fills tcp / rfc5424 / 6.
4. `remove_external_logging` is **destructive — confirm first**: shipping to that destination stops immediately, a compliance/retention gap until reconfigured. Removing the primary promotes the first extra to primary.
## CLI fallback (manifest shape)
There is no `cpln` logging subcommand (`cpln org update --set` covers only description and tags). Get, edit, apply — or `cpln org edit ORG`:
```bash
cpln org get ORG -o yaml-slim > org.yaml # edit spec.logging / spec.extraLogging
cpln apply -f org.yaml --org ORG
```
```yaml
kind: org
name: ORG
spec:
logging:
s3:
bucket: MY_LOG_BUCKET
region: us-east-1
prefix: /
credentials: //secret/AWS_SECRET
extraLogging: # forbidden unless logging is set; max 3
- datadog:
host: http-intake.logs.us3.datadoghq.com
credentials: //secret/DATADOG_KEY
- elastic:
elasticCloud: # variant nests in YAML
cloudId: DEPLOYMENT:BASE64_ID
index: cpln-logs
type: logs # the MCP tool calls this indexType
credentials: //secret/ELASTIC_USERPASS
```
In raw YAML, `syslog` requires all five fields — the documented defaults do not satisfy the required check, and the API rejects an omitted `mode`/`format`/`severity` (the MCP tool fills them for you).
## Template variables
- **CloudWatch** `groupName`/`streamName` accept Fluent Bit record accessors over the shipped fields: `$org`, `$gvc`, `$workload`, `$container`, `$replica`, `$location`, `$provider`, `$version`, `$stream` (e.g. `groupName: $gvc`, `streamName: $workload`). Adjacent variables must be separated by `.` or `,`.
- **Coralogix** `app`/`subsystem` accept only `{org}`, `{gvc}`, `{workload}`, `{location}` — any other `{var}` is rejected at validation.
## What ships
Every entry carries `time` and `log` plus the labels `org`, `gvc`, `workload`, `container`, `replica`, `location`, `provider`, `version`, `stream`. S3 receives gzip-compressed JSONL objects at `PREFIX/ORG/YYYY/MM/DD/HH/MM/UUID.jsonl.gz` (~1 MB chunks). The shipper flushes every 5 seconds; entries appear at the provider within a few minutes.
## Verify
1. `get_external_logging` — primary and extras placed as intended.
2. Generate some traffic, wait 2-5 minutes, check the provider dashboard or bucket.
3. Built-in access is unaffected: `cpln logs '{gvc="GVC", workload="WORKLOAD"}' --org ORG`.
## Troubleshooting
| Symptom | Cause / fix |
|---|---|
| Logs never arrive, no error anywhere | Credential secret has the wrong type — the log router silently skips the provider. Recreate it with the type from the table. |
| S3 stays empty with valid keys | The IAM user lacks `s3:PutObject` on the bucket |
| region/location "must be one of" rejection | CloudWatch and Stackdriver take fixed allowlists, not arbitrary regions |
| `extraLogging` rejected | A primary `spec.logging` must exist first |
| "maximum 3 extra logging providers reached" | 4 providers total is the cap — remove one first |
| xor validation error on a logging block | Exactly one provider key per block — extra providers are separate `extraLogging` entries |
| Datadog/Coralogix/Logz.io value rejected | `host`/`cluster`/`listenerHost` are fixed enums — see the table |
## Quick reference — MCP tools
| Tool | Action |
|---|---|
| `get_external_logging` | Show primary + extra providers |
| `configure_external_logging` | Add or update one provider (automatic primary/extra placement) |
| `remove_external_logging` | Remove a provider (destructive; removing the primary promotes the first extra) |
CLI fallback (CI/CD: `CPLN_TOKEN` + `cpln apply` — read the `cpln` skill first): edit the org manifest as shown above.
## Related skills
| Need | Skill |
|---|---|
| Query logs inside Control Plane (LogQL, Grafana) | `logql-observability` |
| Metrics and tracing export | `metrics-observability` |
| Creating credential secrets, RBAC | `access-control` |
| Org settings: retention, tracing, auth | `org-management` |
## Documentation
- [External Logging Overview](https://docs.controlplane.com/external-logging/overview.md)
- Per provider: [S3](https://docs.controlplane.com/external-logging/s3.md), [CloudWatch](https://docs.controlplane.com/external-logging/cloudwatch.md), [Coralogix](https://docs.controlplane.com/external-logging/coralogix.md), [Datadog](https://docs.controlplane.com/external-logging/datadog.md), [Logz.io](https://docs.controlplane.com/external-logging/logz-io.md), [Stackdriver](https://docs.controlplane.com/external-logging/stackdriver.md), [Syslog](https://docs.controlplane.com/external-logging/syslog.md)
- Elastic, Fluentd, and OpenTelemetry have no docs pages — the MCP tool description is the reference.
firewall-networking10.2 KB
---
name: firewall-networking
description: "Firewall rules and service-to-service communication on Control Plane. Use when the user asks about inbound/outbound rules, CIDR whitelisting, IP blocking, hostname filtering, geo-blocking, header routing, internal endpoints, or network security."
---
# Firewall & Networking
Deep detail for `spec.firewallConfig` and the enforcement model behind it; the `workload` skill owns the summary (deny-by-default, exposure decided at create time, LB picker). Set `firewallConfig` with `create_workload` / `update_workload` — or `public: true`, the shortcut that opens inbound AND outbound to `0.0.0.0/0` (mutually exclusive with an explicit `firewallConfig`). A firewall change creates a new deployment version — a rolling replace, live in about a minute (`vm` workloads are the exception: firewall updates apply in place without restarting the VM).
## How rules are enforced
**Inbound** is checked per request at the mesh sidecar. It counts as fully open only when `inboundAllowCIDR` contains the literal `0.0.0.0/0` AND `inboundBlockedCIDR` is empty; anything else is allow-list mode. Blocked beats allowed; a bare IP means /32. Header and geo filters apply to HTTP traffic only — `tcp`-protocol ports are CIDR-filtered at the connection level instead.
**Outbound** has two separate paths, which is why CIDR rules beat hostname rules:
- **CIDR path** — traffic to `outboundAllowCIDR` ranges bypasses the sidecar and exits directly, on ALL ports unless `outboundAllowPort` is set.
- **Hostname path** — everything else transits the sidecar, which only admits `outboundAllowHostname` entries, matched by Host header (HTTP) or TLS SNI, on ports 80, 443, and 445 (SMB) by default.
- `outboundBlockedCIDR` is subtracted at the network layer and beats both paths — an allowed hostname that resolves into a blocked range still fails.
- Outbound is fully open only with the literal `0.0.0.0/0` in `outboundAllowCIDR`.
## External inbound
```yaml
firewallConfig:
external:
inboundAllowCIDR: # max 250 entries; deduped and sorted on save
- 0.0.0.0/0 # or specific: 203.0.113.0/24, 198.51.100.10
inboundBlockedCIDR: # no max; wins over the allow list
- 192.0.2.0/24
```
## External outbound
```yaml
firewallConfig:
external:
outboundAllowCIDR:
- 198.51.100.0/24 # all ports open to this range while outboundAllowPort is unset
outboundAllowHostname: # lowercase; single wildcard on the prefix only; max 128 chars
- api.stripe.com
- "*.amazonaws.com"
outboundBlockedCIDR:
- 203.0.113.7
```
Source-verified traps:
- **`outboundAllowPort` REPLACES the hostname defaults 80/443/445** — re-list 80 and 443 if you still need them. It also restricts the CIDR path to the listed ports. `protocol` is required (`http`, `https`, or `tcp` — how the proxy treats the port); `number` must be 80 to 65000 and not platform-reserved (8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000).
- **Ports below 80 (22, 25, 53) cannot be listed.** To reach a low port, allow the CIDR and leave `outboundAllowPort` unset — the CIDR path then opens all ports.
- **Private ranges are silently stripped from `outboundAllowCIDR` on managed locations** (10/8, 172.16/12, 192.168/16, 127/8, 169.254/16, 100.64/10, IPv6 ULA): allowing them does nothing, with no error. Reaching a VPC or datacenter takes a wormhole agent (`native-networking`). BYOK clusters keep private ranges.
## Header filters (inbound, HTTP only)
Each filter names a header `key` (max 128 chars) plus exactly ONE of `allowedValues` or `blockedValues` — RE2 regexes; anchor with `^...$` (a bare `bar` also matches `barbell`).
```yaml
firewallConfig:
external:
inboundAllowCIDR: [0.0.0.0/0]
http:
inboundHeaderFilter:
- key: x-api-version
allowedValues: ["^v2$"]
- key: user-agent
blockedValues: ["^BadBot.*", "^Scraper.*"]
```
Matching is OR across everything: a request is rejected if ANY `blockedValues` pattern matches (checked first), and — once at least one allow filter exists — admitted only if ANY `allowedValues` pattern matches. Two allow filters on different headers are alternatives, not both-required; a request missing the header fails its allow filter. **Mesh-internal traffic (10.0.0.0/8 sources) bypasses header filters entirely** — test from outside, not from another workload.
## Geo filtering (country / region / city / ASN)
Two steps: enable geo headers on the workload load balancer (you pick the header names), then filter on those names:
```yaml
spec:
loadBalancer:
geoLocation:
enabled: true
headers: # at least one; names unique; values overwrite client-sent headers
country: x-country
firewallConfig:
external:
inboundAllowCIDR: [0.0.0.0/0]
http:
inboundHeaderFilter:
- key: x-country
allowedValues: ["^US$", "^CA$"]
```
The proxy resolves values from MaxMind GeoLite2 on each request: `country` is the two-letter ISO code (`US`, never `United States`), `region` the subdivision code, `city` the English city name, `asn` the AS number. Echo the headers from the app once before writing filters. HTTP ports only.
## Internal firewall (workload to workload)
`internal.inboundAllowType`: `none` (default), `same-gvc`, `same-org`, or `workload-list`. The admitted identity is the calling workload itself — all its replicas.
```yaml
firewallConfig:
internal:
inboundAllowType: workload-list
inboundAllowWorkload:
- //gvc/GVC/workload/frontend # GVC segment REQUIRED; //workload/NAME is rejected
- /org/ORG/gvc/OTHER-GVC/workload/backend
- cpln://internal/keda # required when a KEDA trigger source is a CP workload
- //agent/DC-AGENT # inbound from behind a wormhole agent (native-networking)
```
- `inboundAllowWorkload` is honored under `same-gvc` too — add specific cross-GVC callers without going `same-org`.
- Links are validated for shape only, never existence — a typo silently denies the caller.
- Internal calls use `http://WORKLOAD.GVC.cpln.local:PORT` (the container port) — plain `http://`, the sidecar adds mTLS. Cross-GVC calls may span locations and then incur egress charges.
## Load balancers (summary)
| Type | Scope | Ports | Static IPs | Wildcard hosts |
|---|---|---|---|---|
| Shared (default) | all workloads | HTTP/HTTPS on 80/443 | no | no |
| Direct | per workload | TCP/UDP, externalPort 22 to 32768 | via IP set | no |
| Dedicated | per GVC (`update_gvc`) | custom domain ports/protocols | via IP set | yes |
```yaml
spec:
loadBalancer:
direct:
enabled: true
ports:
- externalPort: 5432 # 22 to 32768
protocol: TCP # TCP or UDP
containerPort: 5432
```
Direct LB does not terminate TLS (the workload owns its certificates), and its traffic still passes the inbound CIDR rules. `geoLocation` and `replicaDirect` (stateful only) also live under `spec.loadBalancer`. Dedicated LB is a GVC setting (`loadBalancer.dedicated: true`, charged per location) that also carries `trustedProxies` (0 to 2 — which X-Forwarded-For hop counts as the client IP for logging) and a GVC-level `ipSet`. Static IPs and full LB detail: `ipset-load-balancing`.
## Verify
1. `get_resource` (kind="workload") — read `spec.firewallConfig` before changing it, and send the COMPLETE desired `firewallConfig` on update (it replaces as a unit, not field-by-field).
2. `list_deployments` — wait for the new version to report ready in every location.
3. Probe inbound with `curl` from an allowed and a blocked vantage. For an outbound probe from inside the container, use the `cpln` CLI after reading the `cpln` skill.
## Troubleshooting
| Symptom | Cause / fix |
|---|---|
| Outbound to a VPC/private IP fails though its CIDR is allowed | Private ranges are stripped on managed locations — use a wormhole agent (`native-networking`) |
| Hostname egress broke after adding `outboundAllowPort` | The list replaced 80/443/445 — add 80/443 back |
| Need outbound to port 22/25/53 | Below the allowed 80-65000 range — allow the CIDR and leave `outboundAllowPort` unset |
| Header/geo filter not enforced in tests | Testing from another workload (10.0.0.0/8 bypasses header filters), or the port is `tcp` protocol (filters are HTTP-only) |
| Geo allow-list blocks everyone | Values are ISO codes (`^US$`) — full country names never match; echo the header to confirm |
| `workload-list` caller still denied | Link is missing the GVC segment, or has a typo (existence is never validated) |
| KEDA scaler cannot reach its workload trigger source | Add `cpln://internal/keda` to that workload's `inboundAllowWorkload` |
| Firewall seems ignored for one container | `runAsUser: 1337` escapes the mesh and its firewall (see `workload`) |
## Quick reference
| Tool | Purpose |
|---|---|
| `update_workload` | Patch `firewallConfig` (send it complete) or `public` |
| `create_workload` | Decide exposure in the create call: `public: true` or an explicit `firewallConfig` |
| `configure_workload_load_balancer` | Set `spec.loadBalancer` (direct, geo headers, replicaDirect); `remove: true` clears it |
| `update_gvc` | Dedicated LB, `trustedProxies`, GVC-level `ipSet` |
| `get_resource` (kind="workload") / `list_deployments` | Read back config; confirm the rollout |
CLI fallback (no MCP, or CI/CD with `CPLN_TOKEN`): `cpln workload get WORKLOAD --gvc GVC -o yaml > w.yaml`, edit `spec.firewallConfig`, then `cpln apply --file w.yaml --gvc GVC`.
## Related skills
- **workload** — start here: types, spec shape, exposure defaults, internal DNS, LB picker
- **ipset-load-balancing** — static IPs, direct/dedicated LB detail, replicaDirect
- **native-networking** — wormhole agents, PrivateLink/PSC: the answer for private-network traffic
- **cdn-rate-limiting** — CDN in front of workloads, rate limiting
- **workload-security** — JWT authentication, mTLS hardening, direct-LB security
## Documentation
- [Firewall Reference](https://docs.controlplane.com/reference/workload/firewall.md)
- [Load Balancing Reference](https://docs.controlplane.com/reference/workload/load-balancing.md)
- [Service-to-Service Guide](https://docs.controlplane.com/guides/service-to-service.md)
gitops-cicd11.8 KB
---
name: gitops-cicd
description: "Sets up CI/CD pipelines and GitOps for Control Plane. Use when the user asks about GitHub Actions, GitLab CI, Bitbucket, CircleCI, building images in CI, kaniko, cpln apply in pipelines, or service-account tokens for CI."
---
# GitOps & CI/CD
In pipelines the **CLI is the primary interface**: authenticate with a service-account key in `CPLN_TOKEN` (no profile needed), push an image, `cpln apply --ready` the manifests. MCP tools do the work around the pipeline — `get_resource_schema` before authoring manifests, `list_deployments` to confirm a deploy landed. The usual failure is image builds: `cpln image build` runs the build **locally through Docker**, so a runner without a daemon needs a different flow — a daemonless builder, or `--remote` to build on Control Plane. Pick by runner capability, not by habit.
## Service-account authentication
```bash
cpln serviceaccount create --name ci-deployer --org ORG
cpln serviceaccount add-key ci-deployer --description "ci key" --org ORG # --description is required
```
The JSON response's `key` value is the credential — store it as a masked/secret variable in the CI platform. MCP: `add_key_to_service_account` does both steps (and creates the service account if missing).
Grant least privilege (`access-control` skill): pushing images needs `create` on the `image` kind; `cpln apply` needs create/edit on every kind the manifests contain. `cpln group add-member superusers --serviceaccount ci-deployer` works but grants full org access — prefer a scoped policy (`create_policy`).
Set in the platform's variable settings, never inline in scripts:
| Variable | Role |
|---|---|
| `CPLN_TOKEN` | Service-account key (secret/masked) |
| `CPLN_ORG` | Target org |
| `CPLN_GVC` | Target GVC, when the pipeline targets one |
| `CPLN_SKIP_UPDATE_CHECK=1` | Silence CLI update checks in logs |
With `CPLN_TOKEN` set the CLI runs a profile-less session; resolution is flag, then env var, then profile (`cpln` skill). The official example repos persist the token instead — `cpln profile update default --token "$CPLN_TOKEN"` (`create` is an alias of `update`) — either works. Never pass `--token` on ad-hoc commands and never echo the token.
## Installing the CLI on runners
- npm (runner has Node 16+): `npm install -g @controlplane/cli@X.Y.Z` — pin the version. This installs both `cpln` **and** `docker-credential-cpln`.
- Slim or non-Node images: the binary tarball — copy **both** binaries onto PATH. The [containers guide](https://docs.controlplane.com/cli-reference/ci-cd-development/container-image.md) has Dockerfiles for each method, and covers running the CLI inside cron workloads.
## Building images in CI: pick the flow by runner capability
`cpln image build --push` builds locally: with a Dockerfile it shells out to `docker buildx build`, otherwise it downloads the `pack` CLI and runs buildpacks. It needs a working Docker daemon and `docker-credential-cpln` on PATH, and it configures registry auth itself — no separate `docker-login` step. Flag behavior and upload rules: `image` skill.
| Runner | Build flow |
|---|---|
| Daemon available — GitHub-hosted runners, GitLab with the `docker:dind` service (privileged runners, including gitlab.com SaaS), CircleCI `setup_remote_docker`, Bitbucket `docker` service | `cpln image build --name APP:TAG --push`, or keep an existing docker-native pipeline: login below, then `docker build --platform linux/amd64` + `docker push` |
| No daemon — self-managed GitLab runners without privileged mode, locked-down Kubernetes executors | A daemonless builder (kaniko, buildah, rootless BuildKit) pushing straight to the registry, or `cpln image build --name APP:TAG --remote`, which builds on Control Plane with only `CPLN_TOKEN` and the org |
**A remote build gives up the local build options.** `--dockerfile`, `--builder`, `--buildpack`, `--env`, `--env-file`, and `--platform` do not apply — the service detects the build itself and always produces `linux/amd64`. Choose a daemonless builder instead when a job needs build args or a non-amd64 target. `--repo https://github.com/... --branch main` skips the checkout and builds what the service clones; a private repo needs the org's git connection, and in a **non-interactive pipeline the CLI prints an authorization URL and exits**, so authorize it once from a workstation first.
The org registry is a **standard Docker registry**: `ORG.registry.cpln.io`, username = the literal string `<token>`, password = the service-account key. Any tool that can push an OCI image works:
```bash
echo "$CPLN_TOKEN" | docker login ORG.registry.cpln.io -u '<token>' --password-stdin
```
With the CLI installed, `cpln image docker-login` is the faster equivalent for raw `docker push`/`docker pull` jobs: instead of storing a secret it registers the `docker-credential-cpln` helper for the org registry, and Docker resolves the token from `CPLN_TOKEN` (or the profile) at every later call. Use the raw `docker login` form only where the CLI isn't on the box — kaniko auth files, CLI-less build jobs.
GitLab job without a daemon (kaniko; the runner must be amd64 — kaniko cannot cross-build):
```yaml
build:
image:
name: gcr.io/kaniko-project/executor:debug
entrypoint: [""]
script:
- mkdir -p /kaniko/.docker
- printf '{"auths":{"%s.registry.cpln.io":{"username":"<token>","password":"%s"}}}' "$CPLN_ORG" "$CPLN_TOKEN" > /kaniko/.docker/config.json
- /kaniko/executor --context "$CI_PROJECT_DIR" --destination "$CPLN_ORG.registry.cpln.io/my-app:$CI_COMMIT_SHORT_SHA"
```
On GitHub, `docker/login-action` + `docker/build-push-action` also work with the same registry/credentials. Images must be `linux/amd64`; buildpack and multi-platform detail in the `image` skill.
**Tag every build uniquely** (`$CI_COMMIT_SHORT_SHA`, `${GITHUB_SHA:0:7}`). Re-pushing the same tag does not redeploy workloads — if a tag must be reused, set `supportDynamicTags` on the workload or run `cpln workload force-redeployment WORKLOAD` after the push.
## Applying manifests
Author YAML against the real shape first: `get_resource_schema` for each kind. The pipeline then runs:
```bash
cpln apply --file ./manifests/ --ready
```
- `--file` takes a file, a multi-document YAML (`---`), repeated `--file` flags, a directory (recursed; only `.yaml`/`.yml`/`.json` are picked up), or stdin (`--file -`).
- One invocation sorts everything by kind — agent, secret, cloudaccount, gvc, identity, volumeset, policy, workload, then all remaining kinds — so a workload and its GVC can live in one file in any order. `cpln delete --file` applies the reverse order.
- Apply is an upsert. **Renaming a resource in git creates a new resource**; the old one survives until deleted explicitly.
- A manifest with an inline `gvc:` that differs from `--gvc`/`CPLN_GVC` aborts the whole apply.
- `--ready` waits only for the workloads applied in that run: 5-second polls, three consecutive ready checks to pass, a ~5-minute cap, non-zero exit on timeout — a usable deploy gate.
- Seed the repo from a live resource: `cpln workload get NAME -o yaml-slim > workload.yaml` (strips server-managed fields).
- For Helm-chart-shaped releases, `cpln helm install|upgrade|rollback` tracks revisions — the platform's only rollback primitive (`environment-promotion` skill).
## GitHub Actions example
```yaml
name: deploy
on: { push: { branches: [main] } }
env:
CPLN_TOKEN: ${{ secrets.CPLN_TOKEN }}
CPLN_ORG: my-org
CPLN_GVC: my-gvc
jobs:
deploy:
runs-on: ubuntu-latest # GitHub-hosted: Docker daemon available
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 22 }
- run: npm install -g @controlplane/cli@X.Y.Z
- run: cpln image build --name my-app:${GITHUB_SHA:0:7} --push
# manifests reference //image/my-app:IMAGE_TAG — substitute per commit
- run: sed -i "s|IMAGE_TAG|${GITHUB_SHA:0:7}|" manifests/workload.yaml
- run: cpln apply --file ./manifests/ --ready
```
Official starter repos (CLI): [GitHub Actions](https://github.com/controlplane-com/github-actions-example-cli), [GitLab CI](https://gitlab.com/controlplane-com/gitlab-pipeline-example-cli), [Bitbucket](https://bitbucket.org/controlplane-com/bitbucket-pipeline-example-cli), [CircleCI](https://github.com/controlplane-com/circle-ci-pipeline-example-cli), [Google Cloud Build](https://github.com/controlplane-com/google-cloud-build-example-cli). Terraform pipelines: `iac-terraform-pulumi` skill.
## Verify
- In-pipeline: the `cpln apply --ready` exit code is the deploy gate.
- Out-of-band: `list_deployments` for per-location readiness, `get_resource` (kind="image") to confirm the push landed. A workload that never goes ready: `workload` skill or the `/cpln:troubleshoot` command.
## Troubleshooting
| Symptom | Cause / fix |
|---|---|
| `Cannot connect to the Docker daemon` from `cpln image build` | Runner has no daemon — enable dind/privileged mode, switch to a daemonless builder, or build remotely with `--remote` |
| A flag is rejected together with `--remote` | Expected — the service picks the method and always pushes `linux/amd64`, so the local build options do not apply. A job that needs build args or another arch wants a daemonless builder |
| A remote build prints an authorization URL and the job exits | The org has no git connection for that private repo yet — authorize once interactively, then re-run the pipeline |
| Remote build upload rejected — over 500 MB / 20,000 files | Add large paths to `.dockerignore` (or `.gitignore`) — a remote folder build uploads the whole context |
| `The docker-credential-cpln command is not accessible` on `--push` | Helper missing from PATH — npm installs it next to `cpln`; binary installs must copy both binaries |
| `docker login` or push gets 401 | Username must be the literal `<token>`; the key is the password; check it wasn't truncated |
| Push rejected for permissions | Pipeline service account lacks `create` on the `image` kind (`access-control`) |
| `cpln apply` 403 on one kind | Service-account policy doesn't cover that kind — grant per-kind create/edit |
| Apply aborts: `--gvc option ... does not match the gvc value` | Inline `gvc:` in a manifest disagrees with `--gvc`/`CPLN_GVC` |
| Pipeline pushed, workload kept the old code | Same tag re-pushed — use unique tags, `supportDynamicTags`, or `cpln workload force-redeployment` |
| `--ready` exits non-zero after ~5 min | Workload never became ready — check `list_deployments` and workload events |
| `exec format error` at runtime | Image isn't `linux/amd64` (`image` skill) |
## Quick reference
| Tool | Purpose |
|---|---|
| `get_resource_schema` | Manifest shape for any kind before authoring |
| `add_key_to_service_account` | Pipeline service account + key in one call |
| `create_policy` | Scope the pipeline service account's permissions |
| `list_deployments` | Per-location readiness after a deploy |
| `export_terraform` / `convert_to_terraform` | Seed IaC pipelines from live resources or manifests |
## Related skills
| Skill | When |
|---|---|
| `cpln` | CLI conventions — profile-less sessions, flag/env/profile precedence |
| `image` | Build mechanics — the remote vs local table, buildpacks, registry auth, pull secrets |
| `environment-promotion` | Moving images/configs across dev/staging/prod, rollback patterns |
| `iac-terraform-pulumi` | Terraform/Pulumi pipelines instead of `cpln apply` |
| `access-control` | Service accounts, groups, policies, least privilege |
## Documentation
- [CI/CD usage](https://docs.controlplane.com/cli-reference/ci-cd-development/ci-cd.md)
- [Using the CLI in containers](https://docs.controlplane.com/cli-reference/ci-cd-development/container-image.md)
- [CI/CD example repos](https://docs.controlplane.com/guides/gitops.md)
- [cpln apply](https://docs.controlplane.com/guides/cpln-apply.md)
- [Create a service account](https://docs.controlplane.com/guides/create-service-account.md)
iac-terraform-pulumi10.7 KB
---
name: iac-terraform-pulumi
description: "Manages Control Plane resources with Terraform or Pulumi. Use when the user asks about the Terraform or Pulumi provider, infrastructure as code, IaC, exporting resources to HCL, terraform import, state, or drift."
---
# Infrastructure as Code — Terraform & Pulumi
Control Plane has one Terraform provider, `controlplane-com/cpln`. The Pulumi provider (`@pulumiverse/cpln`, published by pulumiverse) is bridged from it, so coverage, semantics, and auth are identical — only the casing changes. The platform also runs a hosted terraform-exporter that converts live resources or schema-validated manifests into provider-correct HCL, reachable through MCP tools and `cpln KIND get -o tf`. The common failure is hand-writing HCL from memory: the nested block shapes are deep and version-specific, and resources that already exist get re-created instead of imported. Generate the HCL, then edit it.
## Choosing an approach
| Approach | Syntax | State | Best for |
|----------|--------|-------|----------|
| Terraform | HCL | Terraform state (use a remote backend) | plan/apply lifecycle, drift detection |
| Pulumi | TypeScript, Python, Go, C# | Pulumi Cloud or self-managed backend | the same lifecycle in a general-purpose language |
| `cpln apply` | YAML/JSON manifests | none — the API is the source of truth | GitOps and CI/CD pipelines (gitops-cicd skill) |
| K8s operator | CRDs | cluster reconcile loop | ArgoCD/Flux shops (k8s-operator skill) |
Pick one owner per resource. A resource managed by Terraform and also edited via console or `cpln apply` shows permanent drift — every `terraform apply` reverts the out-of-band change.
## Provider setup and authentication
```hcl
terraform {
required_providers {
cpln = { source = "controlplane-com/cpln" }
}
}
provider "cpln" {} # configurable entirely via env vars
```
| Provider arg / Pulumi config key | Env var | Notes |
|----------------------------------|---------|-------|
| `org` / `cpln:org` | `CPLN_ORG` | required |
| `token` / `cpln:token` | `CPLN_TOKEN` | service account token for CI/CD |
| `profile` / `cpln:profile` | `CPLN_PROFILE` | local dev: reuse a `cpln login` profile |
| `endpoint` / `cpln:endpoint` | `CPLN_ENDPOINT` | default `https://api.cpln.io` |
| `refresh_token` / `cpln:refreshToken` | `CPLN_REFRESH_TOKEN` | needed only to create an org or update org `auth_config` |
The same env vars drive the `cpln` CLI, Terraform, and Pulumi, so one CI/CD secret serves all three. Create the service account and scope it with a policy (access-control skill); pipeline wiring lives in gitops-cicd.
Pulumi packages: npm `@pulumiverse/cpln`, PyPI `pulumiverse_cpln`, Go `github.com/pulumiverse/pulumi-cpln/sdk/go/cpln`, NuGet `Pulumiverse.Cpln`.
## Coverage
24 resources, all `cpln_` prefixed: agent, audit_context, catalog_template, cloud_account, custom_location, domain, domain_route, group, gvc, helm_release, identity, ipset, location, mk8s, mk8s_kubeconfig, org, org_logging, org_tracing, policy, secret, service_account, service_account_key, volume_set, workload. Data sources: cloud_account, gvc, helm_template, image, images, location, locations, org, secret, workload. Pulumi exposes the same 24 resources in PascalCase (e.g. `CatalogTemplate`).
Per-attribute truth is the registry page for that resource — [Terraform Registry](https://registry.terraform.io/providers/controlplane-com/cpln/latest/docs) or [Pulumi Registry](https://www.pulumi.com/registry/packages/cpln) — not memory. One shape worth knowing up front: `cpln_secret` has no `type` argument; set exactly one per-type attribute (`opaque`, `dictionary`, `aws`, `tls`, ...).
## Generate HCL — don't hand-write it
The hosted terraform-exporter produces provider-correct HCL. Route by what you have:
| You have | Use |
|----------|-----|
| Existing resource(s) | `export_terraform` — a single self link, or bulk by path depth: `/org/ORG` (whole org), `/org/ORG/KIND` (all of a kind), `/org/ORG/gvc/GVC/workload` (all workloads in the GVC) |
| A known set of links | `export_terraform_batch` (full profile) — up to 100 links, merged and de-duplicated; on core, `export_terraform` with path-depth refs covers it |
| A YAML/JSON manifest | `convert_to_terraform` — dry-run validated against the API first, so the returned HCL always matches a schema-valid resource; pass `gvc` for GVC-scoped kinds (workload, identity, volumeset) |
| Nothing yet | author the manifest against `get_resource_schema`, then convert it |
Set `generateImports` on any of these to also get ready-to-run `terraform import` commands, one per resource with the IDs prefilled — run them after `terraform init` and before the first `terraform apply`, so apply updates the live resources instead of re-creating them. `includeDependencies` (export tools) pulls in referenced resources so the HCL is self-contained.
The exporter emits HCL only. For a Pulumi program, convert the exported HCL with the Pulumi CLI: `pulumi convert --from terraform --language typescript --out DIR` (also python, go, csharp, java, yaml). Conversion translates config, not state — adopt the live resources afterwards per Importing below.
The exporter covers 16 kinds: agent, auditctx, cloudaccount, domain (routes emitted as `cpln_domain_route`), group, gvc, identity, ipset, location, mk8s, org, policy, secret, serviceaccount, volumeset, workload. `list_terraform_kinds` (full profile) enumerates them; on core, just attempt the export — an unsupported kind is rejected with the supported list.
**Secrets are never exported.** The upstream exporter would embed revealed values in the HCL, so the MCP tools refuse a ref that targets secrets and refuse wholesale any bulk export that pulled secrets in. Narrow the export to exclude secrets; the user authors secret resources in their own Terraform.
## CLI fallback: -o tf
Without MCP, the same exporter is reachable through the CLI:
```bash
cpln workload get my-app --gvc GVC -o tf > workload.tf # one resource
cpln workload get --gvc GVC -o tf > workloads.tf # no ref: every workload in the GVC
cpln gvc get -o tf > gvcs.tf # every GVC in the org
```
Differences from the MCP tools: the output is bare `resource` blocks only (write the `terraform {}` and `provider "cpln" {}` blocks yourself), no `terraform import` commands, no dependency closure, and no secret guard — never run `-o tf` against a secret (it can print revealed plaintext). Multiple explicit refs in one call are not supported with `-o tf`; export per resource or use the no-ref bulk form. The sibling `-o crd` emits Kubernetes CRD YAML for the operator path (k8s-operator skill). For stateless manifests instead of HCL, use `-o yaml-slim` with `cpln apply`.
## Importing existing resources
Resources must land in state before the first apply, or apply tries to re-create them and fails on name conflicts. `generateImports` returns the exact `terraform import` commands to run — after `terraform init`, before the first apply. Hand-written, the ID is the bare name for org-scoped kinds and `GVC:NAME` for GVC-scoped ones:
```bash
terraform import cpln_gvc.prod prod-gvc
terraform import cpln_workload.api prod-gvc:api
```
Each registry page has an Import section with the exact form (composite kinds differ — a domain route imports as `DOMAIN_LINK:PORT:PREFIX`). On Terraform 1.5+ you may hand-write declarative `import {}` blocks instead; the exporter emits commands, not blocks. Pulumi uses the same IDs (`pulumi import cpln:index/workload:Workload api prod-gvc:api` — the provider is bridged), and `pulumi import --from terraform ./terraform.tfstate` adopts a whole existing Terraform state file into a Pulumi stack.
## Catalog templates
`cpln_catalog_template` (Pulumi `CatalogTemplate`) installs marketplace templates with arguments `name`, `template`, `version`, `gvc`, and `values` (a YAML string). Changing `version` or `values` upgrades the release in place. Template selection and values shapes live in the template-catalog skill.
## Verify
- After an export-and-import, `terraform plan` (or `pulumi preview`) must show zero changes — any diff means the HCL drifted from live state; reconcile before committing.
- Drift detection is the same command on a schedule: a non-empty plan means an out-of-band edit.
- After `terraform apply` on a workload, confirm health with `list_deployments` or `cpln workload get-deployments WORKLOAD --gvc GVC`.
## Troubleshooting
| Symptom | Cause and fix |
|---------|---------------|
| First apply wants to create resources that already exist | The import step was skipped — run the `terraform import` commands from `generateImports` (after `terraform init`), then re-plan |
| `Kind "X" is not Terraform-convertible` | The exporter covers the 16 kinds above (image and user are not among them); manage others via `cpln apply` |
| Export refused mentioning plaintext secrets | Narrow the ref/export to exclude secrets — secret export is not supported |
| Org create or `auth_config` update fails despite a valid token | Those two operations require `CPLN_REFRESH_TOKEN` |
| Pulumi lacks a feature the Terraform provider just shipped | The bridge tracks Terraform provider releases — upgrade the `@pulumiverse/cpln` package version |
## Quick reference
### MCP tools
- `export_terraform` — HCL for existing resources by self link; bulk via path-depth refs; `generateImports`, `includeDependencies`
- `export_terraform_batch` (full profile) — several explicit links merged into one HCL set
- `convert_to_terraform` — manifest to HCL, dry-run validated first
- `list_terraform_kinds` (full profile) — exporter-supported kinds
- `get_resource_schema` — exact API schema when authoring a manifest to convert
CLI fallback: in CI/CD, `CPLN_TOKEN` + `CPLN_ORG` drive `terraform`/`pulumi` directly; `cpln KIND get -o tf` scaffolds HCL from live resources.
### Related skills
| Skill | Use for |
|-------|---------|
| gitops-cicd | pipelines, service account tokens, `cpln apply` workflows |
| k8s-operator | managing resources as Kubernetes CRDs with ArgoCD |
| template-catalog | which template and what values before `cpln_catalog_template` |
| access-control | the service account and policy behind the CI/CD token |
## Documentation
- [IaC Overview](https://docs.controlplane.com/iac/overview.md), [Terraform Provider](https://docs.controlplane.com/iac/terraform.md), [Pulumi Provider](https://docs.controlplane.com/iac/pulumi.md)
- [cpln apply Guide](https://docs.controlplane.com/guides/cpln-apply.md)
- [Terraform examples](https://github.com/controlplane-com/examples/tree/main/terraform); pipeline examples for [GitHub Actions](https://github.com/controlplane-com/github-actions-example-terraform), [GitLab CI](https://gitlab.com/controlplane-com/gitlab-pipeline-example-terraform), and [Bitbucket](https://bitbucket.org/controlplane-com/bitbucket-pipeline-example-terraform)
image16.8 KB
--- name: image description: "Builds, pushes, and manages container images on Control Plane. Use when the user asks about Docker build, remote build, image registry, tags, digests, Dockerfile, buildpacks, pull secrets, ECR/GCR, or image permissions." --- # Control Plane Images Every org gets a private registry at `ORG.registry.cpln.io` — a standard Docker registry (`docker login`/`push`/`pull`/`search` all work). Pushing a tag automatically creates an **image resource** named `NAME:TAG` in the org (read-only `repository`, `tag`, `digest`, `manifest` fields; only metadata tags are editable). An image resource is never created directly — `POST /org/ORG/image` returns 403 "You can create an image only by pushing" — but a **build** can run on Control Plane instead of local Docker (`cpln image build --remote`), which is the answer whenever there is no Docker daemon. The recurring failures are a wrong reference form, a non-`linux/amd64` image (`exec format error`), and a missing or mismatched pull secret. ## Image references | Source | Reference in workload spec | Pull secret | |---|---|---| | Your org's registry (preferred) | `//image/NAME:TAG` (long form `/org/ORG/image/NAME:TAG`) | No — automatic | | Your org's registry (hostname form) | `ORG.registry.cpln.io/NAME:TAG` — the API rewrites it to the link form on write | No — automatic | | Another Control Plane org | `OTHER-ORG.registry.cpln.io/NAME:TAG` — stays a literal external reference | Yes (`docker`) | | Public registry | Exact string: `nginx:latest`, `ghcr.io/owner/app:v1` — **never add a `docker.io/` prefix** | No | | Private external registry (ECR, GCR/GAR, private Docker Hub, ACR...) | Full host path, e.g. `ACCOUNT.dkr.ecr.REGION.amazonaws.com/app:v1` | Yes | - "Image name" always includes the tag (`my-app:v1.0`); a missing tag means `:latest`. Name and tag each max 128 chars; at most two `:` (the second only for a registry port); digest pinning `NAME@sha256:HEX` is supported. - Link references (`/...`) must point into the same org — other orgs are reached only via their registry hostname, which is why they need a pull secret. ## Build and push `--name NAME:TAG` is required — the command aborts without a tag. Everything else follows from **where the build runs**: no Docker daemon on this machine or runner means `--remote`. | | Local (default) | Remote (`--remote`) | |---|---|---| | **Runs on** | Your machine or runner — `docker buildx build` (legacy `docker build` fallback), or `pack` when there is no Dockerfile | Control Plane's build service | | **Needs** | A Docker daemon; `docker-credential-cpln` on PATH for `--push` | The CLI token and org — no daemon, no buildx, no `pack`, no credential helper | | **Source** | `--dir` (default `.`); Dockerfile auto-detected at `DIR/Dockerfile` or given with `--dockerfile PATH` | `--dir` (uploaded) **or** `--repo <https URL>` + optional `--branch` — mutually exclusive | | **Build method** | Dockerfile, else buildpacks (`--builder`/`-B`, `--buildpack`/`-b`; args after `--` go to `pack`) | Auto-detected — Dockerfile when present, otherwise a generated build. Not selectable | | **Push** | Only with `--push` | Always, to the org registry — `--push` is **rejected** | | **Platform** | `--platform`, default `linux/amd64`; a multi-platform list needs buildx **and** `--push` | Always `linux/amd64` — `--platform` is **rejected** | | **Build-time env** | `--env` / `--env-file` — build-time only, **never** runtime | **Not supported yet** — `--env` / `--env-file` are **rejected** with `--remote` | Both need `create` on `image` in the target org. **Rejected with `--remote`:** `--dockerfile`, `--builder`/`-B`, `--buildpack`/`-b`, `--env`, `--env-file`, `--trust-builder`, `--trust-extra-buildpacks`, `--platform`, `--push`. **Remote-only** (they error without `--remote`): `--repo`, `--branch`, `--detach`. ### Remote builds (`--remote`) ```bash cpln image build --name my-app:v1.0 --remote --org my-org # upload and build the current folder cpln image build --name my-app:v1.0 --remote --repo https://github.com/o/a --branch main cpln image build --name my-app:v1.0 --remote --detach # returns a build id, does not watch ``` - **Folder upload** honors `.dockerignore` at the folder root, or `.gitignore` when there is no `.dockerignore`. Caps: 500 MB, 20,000 files, 20,000 directories. `.git`, `node_modules`, `__pycache__`, `.venv`, `venv`, `.mypy_cache`, `.pytest_cache`, `.ruff_cache`, `.DS_Store`, `Thumbs.db`, `desktop.ini` are always excluded (re-include with a `!` rule); `Dockerfile` and `.dockerignore` are force-included. **A symlink pointing outside the folder fails the build.** Rebuilds of the same image NAME upload only changed files — `--no-cache` forces a full re-upload. - **`--repo` takes `github.com` and `gitlab.com` HTTPS URLs only.** An SSH remote or a URL with embedded credentials is rejected; any other host needs support@controlplane.com. - **A private repo builds through the org's git connection, set up once.** The first build that needs it opens a browser to an authorization URL and then continues on its own; in a **non-interactive shell the CLI prints the URL and exits** — authorize, then re-run. The link is single-use and expires shortly. Later builds are silent. - **Watching is not the build.** Logs stream until the push. `Ctrl+C` stops watching and **the build keeps running remotely** — check it with `cpln image get NAME:TAG`. The CLI also gives up watching after 20 minutes, which is not a failure either. ### Local builds ```bash cpln image build --name my-app:v1.0 --push --org my-org # reference as //image/my-app:v1.0 ``` `--push` configures registry auth itself (no separate login step) but hard-errors unless `docker-credential-cpln` is on PATH (installed with the CLI). For a Docker-native flow instead: ```bash cpln image docker-login --org my-org # registers the docker-credential-cpln helper in ~/.docker/config.json docker buildx build --platform=linux/amd64 -t my-org.registry.cpln.io/my-app:v1.0 . docker push my-org.registry.cpln.io/my-app:v1.0 ``` `docker-login` stores no secret — Docker resolves a live token from `CPLN_TOKEN` or the profile on every call. Daemonless CI builders (kaniko, buildah) can also push straight to the registry — it accepts any OCI client with username the **literal string `<token>`** and a service-account key as password (`gitops-cicd` skill). **All images must be `linux/amd64`**; nothing checks the architecture at push or deploy — a wrong-arch image fails at container start with `exec format error`. ### Buildpacks (local builds only) `--remote` detects and builds the source itself and never runs `pack`, so nothing here applies to it. Default builder `heroku/builder:24_linux-amd64`. A `Procfile` (one line: `web: START-COMMAND`) defines the start command; servers must bind `0.0.0.0` and listen on `$PORT`. | Language | Detected by | Trap | |---|---|---| | Node.js | `package.json` + a lockfile | No lockfile means not detected; start from `scripts.start`, `server.js`, or Procfile | | Python, PHP | `requirements.txt` / `uv.lock` / `poetry.lock`; `composer.json` + `composer.lock` | **Procfile REQUIRED** — without it the image builds but exits immediately | | Go, Java, Ruby | `go.mod`; `pom.xml` or `build.gradle` + `gradlew`; `Gemfile` + `Gemfile.lock` | Spring Boot and Rails auto-detected; anything else needs a Procfile | | Rust, C# / .NET | NOT in the default builder | Rust: `-b docker.io/paketocommunity/rust`. .NET: `-B paketobuildpacks/builder-jammy-base` plus `ASPNETCORE_URLS=http://0.0.0.0:$PORT` | ## Pulling - **Same org:** automatic. The platform injects a managed `default-registry` credential into every GVC namespace — never create a pull secret for your own org's images. - **Public images:** no setup. - **Private registries (including other Control Plane orgs):** attach a secret to the **GVC** at `spec.pullSecretLinks` — it applies to all workloads in the GVC; there is no per-workload attachment. Only three secret types work as pull secrets: `docker` (Docker Hub, GHCR, ACR, GAR, other Control Plane orgs — matched to images by registry host in its `auths`), `ecr` (its `repos` list must contain the image's repository; credentials are exchanged for ECR tokens and refreshed automatically), and `gcp` (matched only for images under its own project: `gcr.io/PROJECT/...` or `REGION-docker.pkg.dev/PROJECT/...`). - **Failures are silent:** a linked secret of the wrong type, or one that fails to materialize, is skipped at deploy with no configuration-time error — the symptom is only an image-pull failure on the replica. Attach with `update_gvc` (`pullSecretLinks` is **merged** with existing links; `removePullSecretLinks` removes; an empty list clears all). The registry secret (`docker` — single `dockerConfigJson` string, for another Control Plane org username is the literal `<token>` and the password a service-account key — `ecr`, or `gcp`) must already exist — if it doesn't, offer to draft the manifest for the user to fill and apply (`setup-secret` skill). CLI fallback: `cpln gvc update GVC --set 'spec.pullSecretLinks+=//secret/NAME' --org ORG`. The full cross-org setup (source-org service account, pull policy, target-org secret, GVC) is in the `environment-promotion` skill; `cpln image copy NAME:TAG --to-org ORG2 [--to-name NEW] [--to-profile P] [--cleanup]` is the one-time alternative — it docker-logins both orgs, then pulls, retags, and pushes through the **local Docker daemon** (needs `pull` on the source, `create` on the destination). ## Permissions | Permission | Grants | Implies | |---|---|---| | `create` | Create an image — **this is the push permission** | `pull` | | `pull` | Pull an image (docker pull, cross-org access) | `view` | | `edit` | Modify the image resource — only metadata tags can change | `view` | | `delete` | Delete an image | — | `view` is read-only; `manage` implies all of the above. **Registry authorization is repository-granular.** When the registry checks a docker push or pull it evaluates the permission against `REPOSITORY:*` (tag wildcarded, no resource link) — so a policy whose `targetLinks` list specific `NAME:TAG` images **never authorizes a docker push or pull**. Scope registry policies to all images, or use a `targetQuery` on the `repository` property (`create_policy` with `targetKind: image`; image queries support `name`, `id`, `tag`, `digest`, `repository`, `created`, `lastModified`). Tag-specific `targetLinks` only gate API operations on that image record (view/edit/delete). ## Tags, digests, and redeployment Pushing an existing tag again **updates the same image resource** (new digest, same name) — but running workloads do not follow it by default: - **`supportDynamicTags: false` (default):** the reference is resolved at container start. Serverless pods pull on every start; standard/stateful/cron pods may reuse a node-cached image for any non-`latest` tag — so `cpln workload force-redeployment` after a same-tag push is **not guaranteed** to pick up the new content on every node. - **`supportDynamicTags: true`:** the platform re-resolves every container tag to a digest about every 5 minutes and on each workload change, records the result in `status.resolvedImages` (digest, per-platform manifests, `errorMessages`), and when a digest changes patches the workload — rolling out new pods pinned to `IMAGE@sha256:...`. Digest-pinned references are skipped. For production, prefer immutable tags (commit SHA, semver) or digest pinning; reserve `supportDynamicTags` for dev/staging convenience. ## CLI subcommands `build`, `copy`, `docker-login`, `get [REF...]` (no ref lists all), `delete REF...`, `edit`, `patch`, `query` (`--prop repository=my-app`), `tag` (**metadata** key=value tags on the image resource, NOT docker version tags), `permissions`, `access-report`, `audit`. There is **no `cpln image push` or `pull`** — push via `build --remote`, `build --push`, or `docker push` after `docker-login`. **`copy` has no `--remote` mode and still needs a local Docker daemon.** Verify flags with `cpln image SUBCOMMAND --help` before authoring commands. ## Verify - After a push: `get_resource` (kind="image", name="NAME:TAG") — check `digest` and `lastModified`; `list_resources` (kind="image") to list. - After a workload image change: `list_deployments` for per-location readiness; with dynamic tags, inspect `status.resolvedImages` via `get_resource` (kind="workload") for `errorMessages` and the resolved digest. - After a detached or interrupted remote build: `cpln image get NAME:TAG --org ORG` — the image appears only once the build pushes. `get_image_build` reads a build's status and log by id. - CLI fallback (CI/CD): `CPLN_TOKEN` + `cpln image get NAME:TAG --org ORG -o json`. ## Troubleshooting | Symptom | Cause and fix | |---|---| | `exec format error` at start | Wrong architecture — rebuild with `--platform linux/amd64` | | `docker login` returns "First, grant docker access... cpln image docker-login" | Username was not the literal `<token>` — run `cpln image docker-login`, or login with `-u '<token>'` | | Push/pull 401 "Not authorized to push/pull" | Principal lacks `create` (push) or `pull` on the repository — and tag-scoped `targetLinks` policies never match; use all-images or a `repository` targetQuery | | `cpln image build --push` errors about `docker-credential-cpln` | Helper not on PATH — reinstall the CLI, build with plain docker after `docker-login`, or use `--remote` (which needs no helper) | | `Cannot connect to the Docker daemon` | Only local builds need one — add `--remote` to build on Control Plane instead (`copy` still needs a daemon) | | A flag is rejected together with `--remote` | Expected — the service picks the build method and always pushes `linux/amd64`. Build-time env is not forwarded to a remote build yet; build locally if you need it | | `--repo` / `--branch` / `--detach` errors | Remote-only flags — add `--remote` (and `--branch` also requires `--repo`) | | A remote build prints an authorization URL and exits | Non-interactive shell and the org has no git connection yet — open the URL (single-use, expires shortly), authorize, re-run. Only private repos need it | | Remote build upload rejected — over 500 MB / 20,000 files | Add the large paths to `.dockerignore` (or `.gitignore` if you have no `.dockerignore`) and re-run | | Image pull fails although a pull secret is linked | Wrong secret type (only docker/ecr/gcp work) or host mismatch (`auths` key, ECR `repos` entry, GCP project) — bad secrets are skipped silently | | Same-tag push not picked up | Expected with `supportDynamicTags: false` — see Tags and redeployment above | | `status.resolvedImages.errorMessages`: "unable to parse image" | Resolver limitation for single-segment images with non-alphanumeric tags (`nginx:1.25`) — reference it as `library/nginx:1.25` | | `errorMessages`: "Backing off due to a rate-limit" | Upstream registry returned 429 to tag resolution — wait, or authenticate the registry via a pull secret | | Buildpack image builds but exits immediately | Missing `Procfile` (required for Python and PHP; no web-server auto-detection) | ## Quick reference MCP tools — an image **record** is never created directly (no create-, update-, push-, or copy-image tool); a build is the one write path: - `build_image` — start a build **from a GitHub/GitLab HTTPS repo** and push to the org registry. A build **from a local folder is impossible over MCP** (the server cannot read your filesystem) — route it to `cpln image build --remote --dir PATH`. - `get_image_build` — poll a started build's status and log by the id `build_image` returned. - `list_resources` / `get_resource` (kind="image") — list, or inspect tags/digest/manifest. - `delete_resource` (kind="image", name="NAME:TAG") — removes that image record from the org (destructive). - `update_gvc` — attach existing pull secrets (`docker` / `ecr` / `gcp`, created by the user). - `update_workload` — change a container's image; `get_resource_schema` (kind="image") for the exact resource shape. ### Related skills - **workload** (container spec, where the image reference lives) and **gitops-cicd** (building and pushing from CI) are the usual next hops. - Also: **environment-promotion** (cross-org sharing), **access-control** (policy mechanics), **cpln** (CLI conventions). ## Documentation - [Image Reference](https://docs.controlplane.com/reference/image.md) — resource, permissions, dynamic tags - [Build options — local and remote](https://docs.controlplane.com/cli-reference/get-started/images.md#build-options) — the canonical remote-build reference - [Push an Image](https://docs.controlplane.com/guides/push-image.md) | [Pull an Image](https://docs.controlplane.com/guides/pull-image.md) | [Copy an Image](https://docs.controlplane.com/guides/copy-image.md) - [CLI Image Commands](https://docs.controlplane.com/cli-reference/commands/image.md) · [Buildpacks Guide](https://docs.controlplane.com/guides/buildpacks.md)
ipset-load-balancing10.5 KB
---
name: ipset-load-balancing
description: "Static IPs and load balancers on Control Plane. Use when the user asks about IP sets, fixed IPs, direct or dedicated load balancers, exposing raw TCP/UDP ports, IP allowlisting, geo headers, or egress IPs."
---
# IP Sets & Load Balancing
An IP set reserves one static public IPv4 address per location and attaches it to a **direct** (per-workload) or **dedicated** (per-GVC) load balancer. The linking is bidirectional, and the recurring failure is configuring only one side: the IP set's `spec.link` must point at the workload/GVC AND that target's load balancer must reference the IP set back — otherwise addresses sit `unbound` and the IP set carries `status.warning: Cross-link misconfiguration`. The `workload` skill is primary for the LB-type picker and routing basics; this skill carries the full configuration.
## Load balancer types
| Type | Scope | What it adds | Cost |
|---|---|---|---|
| Default (shared) | every workload | HTTP/HTTPS on 80/443, nothing to configure | included |
| Direct | one workload | raw TCP/UDP on external ports 22-32768, static IPs, TLS passthrough | charged while enabled |
| Dedicated | whole GVC | domain custom ports and TCP routing, wildcard and accept-all hosts, redirects, trusted proxies, static IPs | per location (multiZone adds cross-zone charges) |
Toggling the `direct` block (workload) or `dedicated` flag (GVC) requires the **`configureLoadBalancer` permission** on that resource — `edit` does not imply it (403 "Not allowed to change loadBalancer configuration"); `manage` covers it.
## IP sets
```yaml
kind: ipset
name: partner-ips
spec:
link: //gvc/GVC/workload/WORKLOAD # or //gvc/GVC for a dedicated LB
locations:
- name: //location/aws-us-west-2
retentionPolicy: keep # keep | free
```
How allocation actually works:
- **IPs are allocated in the locations of the linked GVC** (for a workload link, the workload's GVC). No `spec.link`, no allocation — `spec.locations` alone does nothing.
- `spec.locations` pins a per-location `retentionPolicy`; unlisted locations behave as `keep` while in the GVC. Workload links require the GVC segment — `//workload/WORKLOAD` without it is rejected.
- `keep` (default) allocates eagerly and holds the IP through unlinking, GVC location removal, and target deletion (state drops to `unbound`, billing continues until the IP set is deleted). `free` allocates only while bound and releases once the location leaves the GVC or the link/target goes away.
- **Flipping `keep` to `free` does not release an IP whose location is still active in the GVC.** To stop charges: detach the binding (`update_ipset` with `removeLink: true`) so `free` locations release, then delete the IP set to release the rest.
- `state: bound` means both sides point at each other; `unbound` means allocated but unused. Delete is **blocked with 400 while any address is bound** — remove the back-link first. Re-adding a location later does NOT return the same IP.
- Supported on AWS (Elastic IP), GCP (static external address, STANDARD network tier), and Azure (static public IPv4), including BYOK on those clouds. Other providers fail with `status.error` "provider not configured to use IpSets"; cloud IP-quota errors also land in `status.error`.
## Direct load balancer (per workload)
One cloud L4 load balancer per location running the workload, with `externalTrafficPolicy: Local` so the client IP reaches the workload. No TLS termination — the workload owns its certificates. No domain registration needed: each location's address is published on the workload's canonical endpoint DNS with latency-based geo routing, and `status.canonicalEndpoint` switches to the **first** port's `scheme://HOST:externalPort`. Custom hostnames can CNAME to that endpoint. Inbound firewall CIDRs still apply — they become cloud-level source ranges on the LB.
```yaml
spec:
loadBalancer:
direct:
enabled: true
ipSet: //ipset/partner-ips # optional static IPs; that IP set must link back to this workload
ports:
- externalPort: 5432 # 22-32768
protocol: TCP # TCP or UDP
containerPort: 5432 # plain number 80-65535; reserved: 8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000
- externalPort: 443
protocol: TCP
scheme: https # display-only (http|tcp|https|ws|wss): sets the URL scheme shown in UI/status
```
Set with `configure_workload_load_balancer` — it replaces the whole `spec.loadBalancer` block (`remove: true` clears it) and rolls a new deployment (about a minute).
### Geo location headers (`spec.loadBalancer.geoLocation`)
Injects MaxMind GeoLite2 client-location headers on inbound HTTP requests — works with any LB type, no effect on non-HTTP ports. Set `enabled: true` plus `headers` naming at least one of `asn`/`city`/`country`/`region` (names unique, max 128 chars each). Matching client-sent headers are replaced, so apps can trust the values; the country header carries the two-letter ISO code. Filtering on these headers (geo blocking) lives in the `firewall-networking` skill.
### Replica direct (`spec.loadBalancer.replicaDirect: true`)
Stateful workloads only (rejected for other types, including `vm`), capped by a separate quota of **6 replicas per workload**. Each replica becomes addressable as `replica-INDEX.` on the workload's endpoints; internal names appear in `status.replicaInternalNames`. Per-replica custom-domain routing is in the `domain` skill; replica identities and database patterns in `stateful-storage`.
## Dedicated load balancer (per GVC)
A GVC setting — set with `update_gvc` (the `loadBalancer` object is replaced wholesale):
```yaml
spec:
loadBalancer:
dedicated: true
ipSet: //ipset/gvc-ips # optional; that IP set must link back to //gvc/GVC
trustedProxies: 0 # 0 (default) source client IP | 1 last X-Forwarded-For address | 2 second-to-last; sets the logged IP and X-Envoy-External-Address
multiZone: { enabled: false } # cross-zone load balancing, extra charges
redirect:
class:
status5xx: https://errors.example.com # any 500-level response (must be a valid URI)
status401: https://auth.example.com/login?return_to=%REQ(:path)% # supports Envoy format strings
```
Required before domains can use custom ports or the TCP protocol (without it those deploy as warnings and never route) and for wildcard / accept-all hosts — details in the `domain` skill. Enabling or disabling it can cause a brief connectivity blip while DNS propagates. Its access logs are queryable as `{gvc="GVC", workload="_loadbalancer"}`.
## Verify
- `get_resource` (kind="ipset") — every `status.ipAddresses[].state` is `bound`, and no `status.warning` (cross-link) or `status.error` (provider/quota). Share the `ip` values only once bound.
- `list_deployments` — all locations ready after an LB change; the workload's `status.canonicalEndpoint` reflects the direct-LB scheme and port.
- CLI fallback (CI/CD): `CPLN_TOKEN` + `cpln ipset get NAME --org ORG -o yaml`.
## Troubleshooting
| Symptom | Cause and fix |
|---|---|
| IPs stay `unbound`, warning `Cross-link misconfiguration: /org/...` | Only one side is linked — the object named in the warning points here without a matching `spec.link` (or vice versa); configure both sides |
| No IPs allocated at all | `spec.link` missing (locations alone allocate nothing), or the linked GVC has no locations |
| Delete fails 400 "one or more ip addresses are bound" | Remove the workload/GVC back-link or pass `removeLink: true` to `update_ipset`, wait for `unbound`, delete again |
| Still billed after setting `free` | The location is still active in the GVC — `free` releases only when it leaves the GVC or the IP set is unlinked |
| `status.error` "provider not configured to use IpSets" | That location's cloud has no IP-set support (AWS, GCP, Azure only — including BYOK on them) |
| `status.error` AddressLimitExceeded / QUOTA_EXCEEDED / PublicIPCountLimitReached | Cloud-account IP quota exhausted in that region — request an increase from the provider |
| 403 "not granted [configureLoadBalancer]" | Toggling direct/dedicated needs that permission — `edit` alone is not enough |
| API rejects `containerPort` | It is a plain number (80-65535 minus reserved ports); the docs' `containerPort: {port: N}` object form is wrong |
| Deploy warning "TCP access can only be restricted to specific ip addresses when using a custom domain and the GVC has dedicated loadBalancer enabled" | Inbound CIDR rules on a TCP port need the dedicated LB (custom domain) or a direct LB — the shared LB cannot enforce them |
## Quick reference
| Tool | Purpose |
|---|---|
| `create_ipset` | Create with optional `link` and `locations[]` (`retentionPolicy` defaults to `keep`); friendly location names resolve server-side |
| `update_ipset` | Description, tags, replace `link`, or `removeLink: true` to detach |
| `add_ipset_location` | Add locations or overwrite an existing location's `retentionPolicy` |
| `remove_ipset_location` | Drop location entries (releases only IPs whose location is no longer active in the GVC) |
| `list_resources` / `get_resource` / `delete_resource` (kind="ipset") | Read, and delete (releases every IP; blocked while bound) |
| `configure_workload_load_balancer` | Workload side: `direct`, `geoLocation`, `replicaDirect` (`remove: true` clears) |
| `update_gvc` | GVC side: `loadBalancer` (dedicated, ipSet, trustedProxies, multiZone, redirect) |
CLI fallback: `cpln ipset create --name NAME --link LINK --location LOC,POLICY`, plus `add-location` / `update-location` / `remove-location REF --location ...` and `get` / `delete`. `cpln gvc update --set` cannot reach `spec.loadBalancer` — use `cpln gvc edit` or `cpln apply`.
### Related skills
- **workload** — the primary skill: LB-type picker, container ports, endpoints, the `configure_workload_*` tools.
- **domain** — custom domains, custom ports and TCP routes on the dedicated LB, per-replica routing.
- **firewall-networking** — inbound/outbound CIDR rules, header filtering on geo headers.
- **workload-security** — TLS on the workload behind a direct LB, JWT auth, mTLS.
- **stateful-storage** — replica identities and replica-direct with databases.
## Documentation
- [IP Set Reference](https://docs.controlplane.com/reference/ipset.md)
- [Load Balancing Reference](https://docs.controlplane.com/reference/workload/load-balancing.md)
- [GVC Reference (Dedicated LB)](https://docs.controlplane.com/reference/gvc.md)
- [Domain Reference](https://docs.controlplane.com/reference/domain.md)
k8s-operator11.9 KB
---
name: k8s-operator
description: "Manages Control Plane resources as Kubernetes CRDs. Use when the user asks about the Kubernetes operator, kubectl apply for Control Plane, CRDs, ArgoCD GitOps from a cluster, or exporting resources as K8s manifests."
---
# Kubernetes Operator
The operator (Helm chart `cpln-operator`) runs in any Kubernetes cluster and reconciles `cpln.io/v1` custom resources, plus labeled native Secrets, against the platform on a 30-second loop. Reach for it only when resources must live in Git and be reconciled from a cluster (ArgoCD/Flux); for direct provisioning prefer the typed MCP tools, and for pipeline-driven YAML prefer `cpln apply` (`gitops-cicd` skill). The recurring failure is manifest shape: `org`, `gvc`, and `description` sit at the **top level next to `spec`**, not inside it — author the `spec` block with `get_resource_schema`, or skip hand-writing entirely by exporting with `-o crd`.
## Install
cert-manager is a hard requirement (the chart ships a self-signed Issuer whose certificate backs the operator's mutating webhook), then the chart, then per-org auth:
```bash
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.3/cert-manager.yaml
kubectl wait --for=condition=Available deployment --all -n cert-manager --timeout=300s
helm repo add cpln https://controlplane-com.github.io/k8s-operator
helm install cpln-operator cpln/cpln-operator -n controlplane --create-namespace
kubectl get pods -n controlplane -l app=operator
cpln operator install --serviceaccount k8s-operator --org ORG
```
`cpln operator install` configures **auth only** (the Helm step deploys the operator): it gets-or-creates the service account, adds it to a group (`--serviceaccount-group`, default `superusers` — any other group prints a warning), mints a **new key**, and applies a Secret named after the org in the `controlplane` namespace. Re-running with the same `--serviceaccount` is a no-op; a different name replaces the secret with a fresh key; an org-named secret the operator does not own aborts the install. `--export` prints the Secret YAML for Git instead of applying it — but still creates the service account and key on the platform.
Manual equivalent (one secret per org; multiple orgs = multiple secrets):
```yaml
apiVersion: v1
kind: Secret
metadata:
name: ORG # must equal the org name
namespace: controlplane
labels:
app.kubernetes.io/managed-by: cpln-operator # required — the operator only sees labeled Secrets
data:
token: BASE64_KEY # echo -n "KEY" | base64; the CLI also stamps a cpln/serviceaccount annotation it uses to detect reuse
```
Tokens are cached in memory per org and never re-read — after rotating a key, `kubectl rollout restart deployment/operator -n controlplane`. Helm values of note: `env.MANAGE_KINDS` (comma list restricting which kinds get controllers), `env.RECONCILE_INTERVAL_SECONDS` (default 30), `env.CPLN_API_URL`.
## CRD shape
```yaml
apiVersion: cpln.io/v1
kind: workload # kind names are lowercase
metadata:
name: my-app # cluster name; annotation cpln.io/name-replacement overrides the platform name
namespace: default
annotations:
cpln.io/resource-policy: keep # optional: deleting this CR then leaves the platform resource intact
org: ORG # required on every CR — there is no default org
gvc: GVC # required for gvc-scoped kinds: workload, identity, volumeset
description: my app # top level, like tags
spec: # exact platform spec
type: serverless
containers:
- name: main
image: nginx:latest
port: 80
```
Org-scoped kinds: `agent`, `auditctx`, `cloudaccount`, `domain`, `group`, `gvc`, `ipset`, `location`, `org`, `policy`, `serviceaccount`. GVC-scoped: `workload`, `identity`, `volumeset`. Secrets are native v1 Secrets, not CRDs (a `cpln.io/v1 secret` CRD exists but is operator-internal — it holds sync status for native Secrets; never author it). An `mk8scluster` CRD ships but the platform API has no matching endpoint, so mk8s clusters cannot be operator-managed. Recommended layout: one namespace per GVC for gvc-scoped kinds, one per org for org-scoped kinds.
## Secrets (native v1 Secrets)
```yaml
apiVersion: v1
kind: Secret
type: opaque # the platform secret type, lowercase: opaque, aws, azure-connector, azure-sdk, dictionary, docker, ecr, gcp, keypair, nats-account, tls, userpass
metadata:
name: my-secret
namespace: default # any namespace EXCEPT controlplane (everything there is skipped as operator config)
labels:
app.kubernetes.io/managed-by: cpln-operator # required — unlabeled Secrets are invisible to the operator
annotations:
cpln.io/org: ORG # required — selects the org and its auth secret
data: # keys mirror the platform secret's data object
payload: c2VjcmV0LXZhbHVl
encoding: cGxhaW4= # opaque only: plain | base64
```
For `azure-sdk`, `docker`, and `gcp` the platform payload is a single string — put it under one `value` key. Platform tags ride as `cpln.io/`-prefixed annotations. Each synced Secret gets a companion `secrets.cpln.io` CR (same name) carrying sync status and `cpln.io/sync-health-status` / `cpln.io/sync-health-message` annotations — check it when a Secret will not sync.
## How sync behaves
- A mutating webhook stamps every CR (and labeled Secret) with the `cpln.io/sync-protection` finalizer; namespaces labeled `skip-webhook: "true"` are exempt.
- **Local edit (metadata.generation changed):** the operator PUTs the CR to the platform — local state wins.
- **No local edit:** every cycle it pulls platform state into the CR, so console edits appear as CR changes. Under ArgoCD `selfHeal` that registers as drift, Argo restores the Git version, and the operator pushes it back — **Git wins over console edits**.
- **Deleted on the platform:** the CR is deleted from the cluster (with `selfHeal`, Argo re-applies it and the operator re-creates the platform resource).
- **CR deleted:** the platform resource is deleted too, unless annotated `cpln.io/resource-policy: keep`. Deletes blocked by dependent resources retry each cycle — delete children first.
- Failures land in `status.operator.validationError` with the platform error, and retry with exponential backoff (capped at 30s). `status.phase` is `Ready`, `Pending`, `Unhealthy`, or `Suspended`; the platform's own status fields (endpoints, health) are merged into the CR `status`.
- Workloads additionally stream live deployment state over WebSocket into read-only child CRs: `kubectl get deployments.cpln.io` (named `LOCATION.WORKLOAD`), plus `deploymentversions`, `containerstatuses`, and `jobexecutionstatuses` for cron; volumesets get `volumesetstatuslocations` and `persistentvolumestatuses`. Children are owner-referenced and garbage-collected with the parent.
## Exporting existing resources
CRD export is CLI/console-only — no MCP tool emits CRD YAML. Discover and inspect with `list_resources` / `get_resource`, then:
```bash
cpln workload get my-app --gvc GVC --org ORG -o crd > workload.yaml
cpln gvc get -o crd --org ORG > all-gvcs.yaml # no name = whole collection, --- separated
```
In the console, every resource has Export, and the create flow has Preview, with a "K8s CRD" option. System fields are stripped and tags become annotations. **Secret export embeds the revealed payload** (base64-encoded, not encrypted) and needs the `reveal` permission — treat the output as sensitive.
## ArgoCD
Works without special configuration: point an Application at a Git path of CRD manifests (or a Helm chart templating them). The chart patches the ArgoCD ConfigMap with per-kind health checks — CR `status.phase` and `validationError` surface as Argo health — when the `argocd` namespace exists at install time; if Argo came later, run `helm upgrade cpln-operator cpln/cpln-operator`.
```yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata: { name: cpln-resources, namespace: argocd }
spec:
project: default
destination: { server: https://kubernetes.default.svc, namespace: NAMESPACE }
source: { repoURL: "https://github.com/ORG/REPO.git", path: cpln-crds, targetRevision: main }
syncPolicy:
automated:
prune: true # manifest removed from Git = platform resource deleted (honor resource-policy keep)
selfHeal: true # console edits revert to Git
```
## Uninstall
`cpln operator uninstall --org ORG` removes only the auth secret. Before `helm uninstall cpln-operator -n controlplane`, decide the fate of synced resources: annotate CRs `cpln.io/resource-policy: keep` (or delete the CRs you want gone first) — once the operator is gone nothing clears the `cpln.io/sync-protection` finalizer, so leftover CRs stick in Terminating until you strip it (`kubectl patch ... -p '{"metadata":{"finalizers":null}}'`). Platform resources whose CRs were never deleted survive uninstall.
## Verify
- `kubectl get workloads -o yaml` — `status.phase: Ready` and no `status.operator.validationError`.
- `cpln workload get my-app --gvc GVC --org ORG` — the platform side exists and matches.
- `kubectl logs -n controlplane -l app=operator -f` — watch a sync round-trip.
## Troubleshooting
| Symptom | Cause and fix |
|---|---|
| Operator pod not starting, webhook TLS errors | cert-manager missing or certs not issued — `kubectl get pods -n cert-manager`, `kubectl get certificates -n controlplane` |
| Log: "unable to sync resources because the secret ORG could not be found" | Auth secret missing, misnamed, or missing the `managed-by` label (the operator's cache is label-filtered) — re-run `cpln operator install` |
| 401/403 errors after key rotation | The old token is cached — `kubectl rollout restart deployment/operator -n controlplane`; for 403s check the service account's group |
| `validationError`: "CRD resource has no org field" / gvc-scoped kind has no gvc | Add top-level `org` (and `gvc`) — they are not defaulted and do not go in `spec` |
| Secret never syncs, no error anywhere | Missing the label (invisible) or the `cpln.io/org` annotation, or it lives in the `controlplane` namespace (always skipped) — then check the companion `secrets.cpln.io` CR annotations |
| CR stuck Terminating | Platform delete blocked by dependents (delete child resources first) or the operator is gone (strip the `cpln.io/sync-protection` finalizer) |
| Console edits keep reverting | ArgoCD `selfHeal` working as designed — Git is the source of truth; change the manifest instead |
| CRD validation errors on apply | `kubectl explain workload.spec` shows the schema the cluster accepts; regenerate the manifest with `-o crd` |
| `mk8scluster` CR errors with 404 | Expected — the platform API has no mk8scluster path; manage mk8s via `mk8s-byok` instead |
## Quick reference
| Task | Command / tool |
|---|---|
| Author a CRD `spec` block | `get_resource_schema` (kind=workload, gvc, ...) |
| Inspect resources before export | `list_resources` / `get_resource` |
| Configure / remove operator auth | `cpln operator install -s SA --org ORG [--export]` / `cpln operator uninstall --org ORG` |
| Export as CRD manifest | `cpln KIND get [NAME] [--gvc GVC] -o crd`, or console Export / Preview "K8s CRD" |
| Restrict managed kinds | Helm value `env.MANAGE_KINDS: workload,volumeset` |
### Related skills
- **gitops-cicd** — pipelines with `CPLN_TOKEN` + `cpln apply`; choose it over the operator when no cluster-side reconciler is wanted.
- **iac-terraform-pulumi** — the Terraform/Pulumi alternative for declarative management.
- **mk8s-byok** — provisioning a Kubernetes cluster to host the operator (and managing mk8s itself).
- **workload** — the primary skill for what goes inside a workload `spec`.
## Documentation
- [Kubernetes Operator Reference](https://docs.controlplane.com/core/kubernetes-operator.md)
- [Operator Install Guide](https://docs.controlplane.com/guides/cli/cpln-operator.md)
- [Operator source and issues](https://github.com/controlplane-com/k8s-operator)
logql-observability10 KB
---
name: logql-observability
description: "Queries workload logs with LogQL on Control Plane. Use to troubleshoot a workload from its logs, or for log search, access or egress logs, cron run logs, missing or truncated logs, retention, or logs in Grafana."
---
# LogQL & Log Observability
Control Plane stores workload stdout/stderr in Loki and queries it with LogQL. The org is the Loki tenant — it comes from the endpoint path, so `org` is never a query label and queries cannot cross orgs. Reading logs requires the org-level `readLogs` permission, and the in-pod `CPLN_TOKEN` cannot authenticate to the logs endpoint — use a user or service-account token (see the `workload` skill). The recurring agent failure is passing a raw `query` to the MCP tool alongside structured params: a raw query replaces them entirely (the tool rejects the combination), so a raw query must embed every label itself.
## Two ways to query
- **MCP (primary for agents):** `get_workload_logs` — structured params `gvc` (required), `workload`, `container`, `location`, `filter` (literal substring, LogQL `|=`, not regex); window `since` (default `1h`) or `from`/`to` (ISO 8601, `from` inclusive, `to` exclusive); `limit` (default 30, max 999, single request — `truncated: true` means narrow the window or filter harder); `order` (`oldest_first` default, or `newest_first`). For regex, parsers, or the `replica`/`stream`/`version` labels, pass a raw `query` (max 500 chars) instead of the structured selectors.
- **CLI:** `cpln logs '<LOGQL>'` — interactive debugging (live `--tail`), scripts, CI.
```bash
# Defaults: --since 1h, --limit 30, --direction forward
cpln logs '{gvc="GVC", workload="WORKLOAD"}' --org ORG
cpln logs '{gvc="GVC", workload="WORKLOAD"} |= "error"' --limit 100
cpln logs '{gvc="GVC", workload="WORKLOAD", container="main"}' --since 7d
cpln logs '{gvc="GVC", workload="WORKLOAD"}' --tail # live follow; the server ends a tail session after 30m
cpln logs '{gvc="GVC", workload="WORKLOAD"}' \
--from 2026-06-01T00:00:00Z --to 2026-06-02T00:00:00Z # ISO 8601 or relative (7d, now-1M); from inclusive, to exclusive
cpln logs '{gvc="GVC", workload="WORKLOAD"} |= "error"' --since 24h --limit 0 # 0 = unlimited, auto-paginates
cpln logs '{gvc="GVC", workload="WORKLOAD"}' -o jsonl # one JSON object per line; -o raw = bare lines
```
- `--since` takes relative durations (`ms s m h d w mo y`, compound like `1h30m`). `--from`/`--to` (CLI) take an ISO 8601 timestamp **or** a relative duration meaning that long ago — `--from 7d`, `--from now-1M`, `--to now-30m` (the CLI accepts `M` for months; the `get_workload_logs` tool takes ISO 8601 only).
- `cpln workload eventlog WORKLOAD` (alias `cpln workload log`) is resource event history, not container output — for application logs always use `cpln logs`.
## Labels
| Label | Value |
|:---|:---|
| `gvc` | GVC name |
| `workload` | Workload name |
| `container` | Container name, or a built-in stream below |
| `location` | Deployment location, e.g. `aws-us-east-1` |
| `provider` | Cloud provider |
| `replica` | Replica (pod) name — unique per cron execution |
| `stream` | `stdout` or `stderr` |
| `version` | Workload deployment version that wrote the line |
At least one non-empty matcher is required; regex matchers work — `{gvc=~".+"}` spans every GVC in the org.
## Filters and LogQL features
| Operator | Meaning | Example |
|:---|:---|:---|
| `\|= "text"` | contains | `\|= "error"` |
| `!= "text"` | does not contain | `!= "health"` |
| `\|~ "regex"` | matches regex | `\|~ "timeout\|crash"` |
| `!~ "regex"` | does not match | `!~ "debug\|trace"` |
Loki is current (3.x), so full LogQL works: parsers (`| json`, `| logfmt`, `| pattern`, `| regexp`), post-parse label filters (`| latency > 100`), `line_format`, and metric queries (`count_over_time`, `rate`, `sum ... by`). The CLI and MCP tool print log lines only — run metric queries in Grafana: the `Explore on Grafana` link on the console Logs page opens it with the query prefilled (the org `grafanaAdmin` permission grants the Grafana Admin role; everyone else is Viewer).
```logql
{gvc="GVC", workload="WORKLOAD"} |= "error" != "health" # errors minus noise
{gvc="GVC", workload="WORKLOAD"} |~ "panic|fatal|exception" # crashes and stack traces
{gvc="GVC", workload="WORKLOAD", container="_accesslog"} |= "\" 50" # HTTP 5xx in access logs
sum(count_over_time({gvc="GVC", container="_accesslog"} |= "\" 50"[1m])) by (workload) # 5xx rate (Grafana)
```
## Built-in log streams
| Selector | Contents |
|:---|:---|
| `container="_accesslog"` | Inbound requests (Envoy access-log format) on the workload's ports |
| `container="_requestlog"` | The workload's outbound (egress) requests through the sidecar, same format |
| `workload="_loadbalancer"` | Access logs of the GVC's dedicated load balancer |
| `container="_alerts"` | Threat-detection (Falco) alerts, with extra labels `rule`, `priority`, `source` |
Platform health probes and unroutable-request noise are filtered out of `_accesslog` by design, so probe traffic never shows up there. Sidecar and system containers (`istio-init`, `istio-validation`, `cpln-*`, `debugger-*`) are never collected.
## Why logs go missing (pipeline limits)
- **Lines over 16 KiB are cut** at 16 KiB with a `... [truncated N bytes]` suffix; empty lines are dropped.
- **Per-replica rate limit:** each replica+container pair is capped (roughly 10,000 lines/s on managed clusters; effectively unlimited on BYOK). Excess lines are dropped, and when collection resumes one marker line appears: `#### Replica logs were rate-limited: N lines in the last Xs were not collected ####`.
- **Retention:** org spec `observability.logsRetentionDays`, default 30 (0 turns log collection off entirely; the `org-management` skill owns org spec edits). Queries beyond retention return nothing, not an error.
- **Server caps:** one query may span at most 31 days and times out after 2 minutes; tail sessions end after 30 minutes.
## Cron workloads: logs for one execution
`{gvc=, workload=}` on a cron workload interleaves every past run. Each execution runs in its own replica, so scope to one run with the `replica` label plus the execution's time window.
**1. List executions.** `list_deployments` **with** `location` returns the full deployment JSON; `status.jobExecutions[]` holds the per-execution metadata (without `location` the tool returns only a readiness summary). CLI:
```bash
cpln workload get-deployments WORKLOAD --gvc GVC -o json | jq '
[(.items // .)[] as $d | $d.status.jobExecutions[]? | . + {location: $d.name}]
| sort_by(.startTime)[] | {location, name, status, startTime, completionTime, replica}'
```
Schema facts (nodelibs `cronjob.ts`) that bite scripts:
- `status` is one of `successful | failed | active | pending | invalid | removed` (default `pending`). Running means `status == "active"` AND no `completionTime`; a missing `completionTime` alone proves nothing (the schema warns it is not an indication of success).
- `completionTime: null` is stripped server-side and `startTime` may be absent for never-started runs — each field is present-with-value or absent, so guard flags with `${VAR:+--from "$VAR"}` and jq `// empty`.
- `replica` is optional: when missing, the run never got a pod and there are **no logs** — diagnose with the execution's `containers` map and `message` field (aggregated pod events) and `get_workload_events`.
**2. Query that replica, time-bounded.** MCP: raw `query` embedding ALL labels (structured params must be omitted) plus `from`/`to`. CLI:
```bash
cpln logs '{gvc="GVC", workload="WORKLOAD", location="LOCATION", replica="REPLICA-ID"}' \
--org ORG --from START_TIME --to COMPLETION_TIME
```
`jobExecutions` timestamps are ISO 8601 and pass straight through. Pad the window a minute or two on each side, and widen it before concluding logs do not exist. For a live run (`active`, no `completionTime`), drop `--to` and add `--tail`.
## Quick reference
| Tool | Use |
|:---|:---|
| `get_workload_logs` | LogQL queries — structured params or raw `query` |
| `list_deployments` | Deployment health; cron `status.jobExecutions` (pass `location`) |
| `get_workload_events` | Probe failures, scheduling, restarts — events, not app logs |
CI/CD and headless use: set `CPLN_TOKEN` and run `cpln logs` directly (no profile needed); the principal must hold org `readLogs`.
## Troubleshooting
| Symptom | Cause and fix |
|:---|:---|
| 403 `requires permission "readLogs" in org` | Grant `readLogs` on the org via policy (it implies `view`) |
| `Invalid --from format: ...` | `--from`/`--to` accept an ISO 8601 timestamp or a relative duration (`7d`, `now-1M`); correct the value |
| Empty result for known activity | Window outside retention (default 30d), span over 31 days, or filters too narrow; app may not write to stdout/stderr |
| `#### Replica logs were rate-limited ... ####` marker | Per-replica cap was hit; lines in that interval are gone — reduce log volume |
| Line ends with `... [truncated N bytes]` | 16 KiB per-line cap — emit smaller lines |
| Health checks absent from `_accesslog` | Filtered by design; probe failures surface in `get_workload_events` |
| MCP rejects raw `query` combined with `workload` etc. | A raw query replaces the structured params — embed all labels in the query itself |
| `cpln workload log` shows no app output | That is the eventlog alias; use `cpln logs` |
## Related skills
| Skill | Owns |
|:---|:---|
| `workload` | Deploy and diagnose flow, injected `CPLN_*` env vars, canonical URLs |
| `metrics-observability` | PromQL, default metrics, Grafana alert rules, Prometheus federation |
| `external-logging` | Shipping logs to S3, Datadog, Coralogix, and other providers |
| `mk8s-byok` | mk8s cluster logs add-on (`cluster_name` / `namespace` labels) |
## Documentation
- [Logs Reference](https://docs.controlplane.com/core/logs.md)
- [CLI logs Command](https://docs.controlplane.com/cli-reference/commands/logs.md)
- [External Logging Overview](https://docs.controlplane.com/external-logging/overview.md)
- [LogQL (upstream Grafana reference)](https://grafana.com/docs/loki/latest/query/)
metrics-observability14.1 KB
---
name: metrics-observability
description: "Workload metrics, PromQL, Grafana, and tracing on Control Plane. Use to observe or troubleshoot a workload via CPU/memory/request or custom metrics, traces, Prometheus federation, alerts, or metrics retention."
---
# Metrics, Tracing & Observability
Control Plane stores every workload's metrics as Prometheus-compatible time series in a managed backend (Mimir), queryable in PromQL through the per-org managed Grafana or the MCP tools. The org is the tenant — it comes from the endpoint path, so there is no `org=` label and no cross-org queries. Two traps dominate. **Series names are short:** memory is `mem_used` / `mem_reserved` / `mem_billable`, not `memory_*` — a `memory_used` query returns nothing, so ground names with `list_metrics` first. **Rate-shaped metrics are pre-rated:** `egress`, `requests_per_second`, and the latency buckets are already rated by the platform's recording rules, so you query them bare — wrapping them in `rate()` again returns garbage. Finally, a workload's in-pod `CPLN_TOKEN` cannot authenticate to the metrics endpoint; querying from outside the mesh needs a user or service-account token.
## Two ways to query
- **MCP (primary for agents):** `query_metrics` runs a PromQL query — a range query over the last `1h` at `60s` step by default; pass `resolution: "instant"` for a single point, or `since` / `from` / `to` / `step` to adjust. `list_metrics` discovers the metric names and real label values present in the org right now (built-in, `kube_`/`node_`, and custom); pass `metric:` to ground one metric's live labels before filtering. Reach for it whenever a query returns no series. Measure first, then change scaling settings.
- **Grafana:** the managed per-org instance — open **Metrics** in the Console sidebar (or the **Metrics** link on any workload), use **Explore** for ad-hoc PromQL, and dashboards/alerting for the rest. The `grafanaAdmin` org permission grants the Grafana Admin role; everyone else is Viewer.
`list_metrics`' built-in catalog still spells memory `memory_*`; trust the live names it returns (and this skill) — the queryable series is `mem_*`.
## PromQL: query the right shape
The platform pre-computes rates, so the shape decides the query form:
- **Gauges — query bare:** `cpu_used`, `mem_used`, `replica_count`, `workload_ready_replicas`.
- **Pre-rated gauges — query bare, never `rate()`:** `egress` and `cross_zone_traffic` (bytes per minute), `requests_per_second`, `requests_initiated_per_second`, `cron_execution_rate`.
- **Histogram — `histogram_quantile`, no extra `rate()`:** `request_duration_ms_bucket` keeps its `le` label and is already rated.
- **Cumulative counters — wrap in `increase()` / `rate()` for velocity:** `container_restarts`, `cron_executions`, `workload_progress_failure`, `workload_rescheduled_replicas`, `domain_warnings`.
```promql
cpu_used # cores in use, per replica (bare gauge)
sum by (workload) (mem_used) # memory bytes per workload — mem_, not memory_
egress # outbound bytes/minute (already rated — no rate())
sum by (workload) (requests_per_second{response_class="500"}) # 5xx rate; response_class is "200".."500"
histogram_quantile(0.95, sum by (le) (request_duration_ms_bucket)) # p95 latency (ms); no rate() wrapper
sum by (gvc, workload) (increase(container_restarts[5m])) # restarts in the last 5m
```
## Built-in metrics
Collected for every workload, no configuration. Names and types below are the recording-rule outputs (the queryable series). Call `list_metrics` for the complete live set, including your custom metrics.
**Resource & network** (per replica): `cpu_used` / `cpu_reserved` / `cpu_billable` (cores, gauge); `mem_used` / `mem_reserved` / `mem_billable` (bytes, gauge); `egress` / `cross_zone_traffic` (bytes/minute, pre-rated gauge); `replica_count` / `workload_ready_replicas` / `workload_desired_replicas` (gauge).
**Traffic** (per pod): `requests_per_second` and `requests_initiated_per_second` (pre-rated gauge, label `response_class`); `request_duration_ms_bucket` (latency histogram, keeps `le`).
**Stability**: `container_restarts`, `workload_progress_failure`, `workload_rescheduled_replicas`, `cron_executions`, `domain_warnings` (cumulative counters); `cron_execution_rate` (pre-rated); `capacity_ai_updates`, `load_balancer` (gauge).
**Volume** (per volume set): `volume_set_capacity_bytes`, `volume_set_used_bytes`, `volume_set_free_bytes`, `volume_set_billable_bytes`, `volume_set_capacity_billable`, `volume_set_snapshots_billable`.
**Org-wide** (no workload label): `logs_storage_mb` / `metrics_storage_mb` / `tracing_storage_mb`; `agent_peers_count` / `agent_services_count` (gauge) and `agent_{tx,rx}_{bytes,packets}_total` (counter) from wormhole agents; `threat_detection_alerts` / `threat_detection_forward_total` / `threat_detection_forward_enabled`.
mk8s clusters with metrics enabled also expose `kube_*` (kube-state-metrics) and `node_*` (node-exporter).
## Custom metrics
A container exposes Prometheus-format metrics by declaring a `metrics` block; the platform scrapes every replica every **30 seconds** (5s timeout). Set it at creation with `create_workload` or add it later with `update_workload`; if the typed tool doesn't surface the nested field, fall back to `get_resource_schema` for `workload` then `cpln apply -f workload.yaml`.
```yaml
spec:
containers:
- name: app
metrics:
port: 9100 # required; ≥80 and NOT a reserved port (see trap below)
path: /metrics # required; string, max 128, default /metrics
dropMetrics: # optional; RE2 regexes, dropped before scrape
- '^go_.*'
- '^process_.*'
```
- **Reserved-port trap:** `port` rejects the platform's sidecar ports — `9090`, `9091`, `8012`, `8022`, `15000`/`15001`/`15006`/`15020`/`15021`/`15090`, `41000`. The obvious Prometheus default `9090` fails; use `9100`, `2112`, etc.
- Metric names starting with `cpln_` are dropped (you cannot overwrite platform series).
- Scraped samples gain labels `org`, `gvc`, `workload`, `container`, `location`, `provider`, `region`, `cluster_id`, `replica`.
## Distributed tracing
Tracing answers a different question than metrics: not "is latency high?" but **where** in the request path. It is opt-in via `spec.tracing` on a **GVC** (or org-wide on the org spec) — set it with `update_gvc` / `create_gvc` or `cpln apply`. Exactly one provider (`.xor`), and `sampling` (a required `0`–`100` percentage):
- **`controlplane`** — built-in backend, queryable with the tools below; zero extra infrastructure.
- **`otel`** — ship spans to your own OpenTelemetry collector (`endpoint`).
- **`lightstep`** — ship to Lightstep (`endpoint` + an opaque `credentials` secret).
`customTags` adds fixed key/values to every span (each value max 50 chars). Only requests served after enablement, in the sampled fraction, produce traces. Apps wanting to emit their own spans to the `controlplane` provider send OTLP to `tracing.controlplane:80` (gRPC) or `tracing.controlplane:4318` (HTTP).
Query the built-in backend with `query_traces` — structured params (`gvc`, `workload`, `location`, `errorsOnly`, `minDuration: "500ms"`) or a raw `traceql` query that replaces them; span attributes are `resource.gvc` / `resource.workload` / `resource.location`. Then `get_trace` reads one trace's span tree to name the slow/failed span. **Empty results are usually configuration:** confirm tracing is enabled, sampling catches traffic, and the window saw requests. Triage flow: `query_traces` (`minDuration` or `errorsOnly`) to the worst trace, `get_trace` to the culprit span, then `get_workload_logs` over the same window for the application error.
## Built-in Grafana alert rules
The managed Grafana provisions five rules, all annotated to the `cpln-metrics-overview` dashboard. They evaluate on import but deliver nothing until a Grafana contact point exists — set `defaultAlertEmails` (below) or add a contact point. Deletions are recreated on next login.
| Rule | Fires when | Default |
|:---|:---|:---|
| `container-restarts` | `increase(container_restarts[5m]) > 0` per gvc/location/workload (any restart) | active |
| `stuck-deployments` | more than one deploy `version` of a workload is restarting (15m) | active |
| `workload-progress-failure` | `increase(workload_progress_failure[10m]) > 0` (15m) | active |
| `threat-detection-alerts` | `increase(threat_detection_alerts[15m]) > 0` per gvc/workload/priority/rule | active |
| `domain-warnings` | `increase(domain_warnings[60m]) > 5` per domain/type | **paused** |
## Retention & billing
Retention and default alert recipients live in the org `observability` block. No typed MCP tool edits it (the `org-management` skill owns org-spec edits) — apply via CLI: `get_resource_schema` for `org`, then `cpln apply -f org.yaml`.
```yaml
kind: org
spec:
observability:
logsRetentionDays: 30 # int 0-3650, default 30 (0 disables log collection)
metricsRetentionDays: 30 # int 0-3650, default 30
tracesRetentionDays: 30 # int 0-3650, default 30
defaultAlertEmails: # email[]; recipients for the grafana-default-email contact point
- ops@example.com
```
Combined storage of logs, metrics, and traces is charged per GB-month over 100 GB.
## Export & centralize metrics
`readMetrics` (org permission, "access usage and performance metrics") gates both the federation endpoint and Grafana data sources. Create a service account and grant it `readMetrics` via policy — the `access-control` skill owns that; here is the metrics-specific wiring.
**Federate into your own Prometheus** — scrape the source org with the service-account token:
```yaml
scrape_configs:
- job_name: cpln-federate
scheme: https
honor_labels: true
metrics_path: '/metrics/org/SOURCE_ORG/api/v1/federate'
params:
'match[]': ['{__name__=~".+"}'] # narrow the matcher to limit egress
authorization: { type: Bearer, credentials: "${CPLN_SERVICE_ACCOUNT_TOKEN}" }
static_configs:
- targets: ['metrics.cpln.io']
```
**Cross-org Grafana** — in a viewer org's Grafana, add a Prometheus data source with URL `https://metrics.cpln.io/metrics/org/SOURCE_ORG` and a custom HTTP header `authorization` = `Bearer <SOURCE_ORG_SA_TOKEN>`, then **Save & Test**. The community dashboard `grafana.com/dashboards/20378` (Multi-Source Metrics Overview) visualizes several at once.
**Token trap:** `metrics.cpln.io` authenticates user and service-account tokens only. A workload's injected `CPLN_TOKEN` does **not** work there even with `readMetrics` on its identity — the metrics proxy forwards only the link headers, never the signed header the in-mesh API path injects (see the `workload` skill). Query from inside a workload with a service-account key.
## Autoscaling metric availability
This skill covers only which scaling metrics each workload type allows; for strategy, YAML, percentiles, multi-metric, KEDA, and Capacity AI, see the `autoscaling-capacity` skill.
| Metric | Serverless | Standard | Stateful |
|:---|:---:|:---:|:---:|
| `concurrency` | yes | no | no |
| `cpu` / `memory` / `rps` | yes | yes | yes |
| `latency` / `keda` | no | yes | yes |
| `disabled` | yes | yes | yes |
`vm` workloads allow only `disabled`; `cron` has no autoscaling. (`memory` here is the scaling keyword — distinct from the `mem_used` series.)
## Quick reference
| Tool | Use |
|:---|:---|
| `list_metrics` | Discover real metric names and label values (built-in + custom) before querying |
| `query_metrics` | Run a PromQL query against the org's metrics |
| `query_traces` | Search traces (TraceQL) — slow (`minDuration`) or failed (`errorsOnly`) requests |
| `get_trace` | Read one trace's span tree to locate the slow/failed span |
| `get_workload_logs` | Correlate a metric spike with logs (see `logql-observability`) |
- **Metrics endpoint:** `https://metrics.cpln.io/metrics/org/{ORG}` (federation adds `/api/v1/federate`).
- **Permission:** `readMetrics` (federation endpoint + Grafana data source).
- No typed tool edits the org `observability` block, the GVC `tracing` block, or a container `metrics` block — fall back to `get_resource_schema` + `cpln apply`.
## Troubleshooting
| Symptom | Cause and fix |
|:---|:---|
| Query returns no series | Wrong name — memory is `mem_used`, not `memory_used`; run `list_metrics` to confirm live names |
| `egress`/latency values look tiny or wrong | Pre-rated series wrapped in `rate()` — query `egress` bare, latency via `histogram_quantile(..., request_duration_ms_bucket)` |
| Custom `metrics` block rejected | `port` is reserved (`9090`/`9091`/`15000`+) or below 80 — use `9100`/`2112` |
| Custom metrics never appear | Names prefixed `cpln_` are dropped; scrape runs every 30s — allow a cycle |
| 403 at `metrics.cpln.io` | Principal lacks `readMetrics`, or an in-pod `CPLN_TOKEN` was used — use a user/SA token |
| `query_traces` empty | Tracing not enabled on the GVC, sampling too low, or no traffic in the window |
| Alert never notifies | Rules evaluate but need a contact point — set `defaultAlertEmails` or add one in Grafana (`domain-warnings` also ships paused) |
## Related skills
| Skill | Owns |
|:---|:---|
| `workload` | Deploy/diagnose flow, injected `CPLN_*` env vars, the spec that holds `metrics` |
| `autoscaling-capacity` | Scaling strategy, per-metric YAML, percentiles, KEDA, Capacity AI |
| `logql-observability` | Log queries (LogQL), `cpln logs`, correlating spikes with log events |
| `org-management` | Org-spec edits — the `observability` retention block |
| `external-logging` | Shipping logs to S3, Datadog, Coralogix, and other providers |
## Documentation
- [Default Metrics](https://docs.controlplane.com/guides/default-metrics.md)
- [Custom Metrics](https://docs.controlplane.com/reference/workload/custom-metrics.md)
- [Export Metrics (federation)](https://docs.controlplane.com/guides/export-metrics.md)
- [Centralized Metrics](https://docs.controlplane.com/guides/centralized-metrics-management.md)
- [Autoscaling](https://docs.controlplane.com/reference/workload/autoscaling.md)
- [PromQL (upstream Prometheus reference)](https://prometheus.io/docs/prometheus/latest/querying/basics/)
migration-patterns14.8 KB
---
name: migration-patterns
description: "Migrate workloads from Kubernetes, Docker Compose, or Helm to Control Plane. Use when the user asks to convert k8s manifests, a docker-compose.yml, or Helm charts, or to move an existing app onto the platform."
---
# Migrating to Control Plane
Each source format has its own converter, and they are not interchangeable: Kubernetes through `cpln convert`, Docker Compose through `cpln stack`, a Helm chart of Control Plane resources through `cpln helm`. All three are **CLI-only — there is no MCP converter.** The dominant failure is hand-translating a Compose/k8s/Helm artifact into Control Plane YAML — even one "small enough to do by hand" — instead of running the tool and then reviewing what it left behind. The converter gets the mechanical translation right; your value is the gap analysis on top of it. If asked to translate by hand, push back: convert first, then work through the fix-ups.
## Pick the conversion path
| Source | Convert (CLI-only) | One-shot deploy |
|---|---|---|
| Kubernetes manifests | `cpln convert -f k8s.yaml --gvc GVC` | `cpln apply -f k8s.yaml --k8s true` |
| Kubernetes Helm chart | `helm template R ./chart \| cpln convert -f - --gvc GVC` | — |
| Docker Compose | `cpln stack manifest --gvc GVC` (preview) | `cpln stack deploy --gvc GVC` |
| Helm chart of CPLN resources | — | `cpln helm install R ./chart --gvc GVC` |
`cpln helm` does **not** convert Kubernetes manifests — its charts must render only Control Plane kinds. There is no `cpln stack convert`; `cpln stack manifest` previews the generated YAML.
## Kubernetes (`cpln convert`)
`cpln convert -f FILE [--gvc GVC]` reads a single file, a directory (recursive), or `-` for stdin, and writes Control Plane YAML. `cpln apply -f FILE --k8s true` runs the same converter and applies the result in one step. Resources that become their own Control Plane resource:
| K8s resource | Control Plane resource |
|---|---|
| Deployment, StatefulSet, ReplicaSet, ReplicationController, DaemonSet | Workload (type derived) |
| CronJob | Workload (cron, schedule from spec) |
| Job | Workload (cron, default schedule `* * * * *`) |
| Secret | Secret (type-mapped, below) |
| ConfigMap | Secret (dictionary) |
| Ingress | Domain (with routes) |
| PersistentVolumeClaim | VolumeSet |
Resources that shape the conversion without becoming their own resource: **Service** (port-protocol inference, public-exposure detection that sets the workload firewall, ingress route resolution), **HorizontalPodAutoscaler** (workload `minScale`/`maxScale`/`scaleToZeroDelay`/CPU target), **ServiceAccount** (image pull-secret extraction), **PersistentVolume + StorageClass** (volumeset capacity, performance class, filesystem), **EndpointSlice** (pod mapping for selectorless services).
**Workload type — `cron > stateful > standard`:** a Job/CronJob becomes `cron`; otherwise any container mounting a volumeset (from a PVC or `volumeClaimTemplates`) becomes `stateful`; everything else is `standard`. The converter never emits `serverless` or `vm` — switch a workload to those yourself after converting.
**Secret type mapping:**
| K8s secret | Control Plane type |
|---|---|
| `kubernetes.io/dockerconfigjson` | `docker` |
| any data key named `payload` | `opaque` |
| `kubernetes.io/basic-auth` | `userpass` |
| `kubernetes.io/tls` | `dictionary` (validated for `tls.crt`/`tls.key`, stored as a dictionary) |
| everything else / ConfigMap | `dictionary` |
**PVC performance class:** `io1`, `io2`, `pd-extreme`, `UltraSSD_LRS`, `thick`, `fast`, `persistent_1` map to `high-throughput-ssd` (matched on the StorageClass parameter value); everything else — `gp2`, `gp3`, the default — maps to `general-purpose-ssd`.
**Port protocol** (when `--protocol` is not forced): Service `appProtocol` wins outright; otherwise the converter gathers hints from the Service and container port-name prefixes, the probe type, and the port number, then picks the most specific (grpc > http2 > http > tcp); default `tcp`.
The converter auto-creates an identity `identity-<workload>` and policy `policy-<workload>` granting `reveal` for every workload that references secrets. When `--gvc` is omitted, workload links carry a `{{GVC}}` placeholder — replace it before applying.
## What `cpln convert` leaves for you
The converter translates structure faithfully but **warns on only two things** — a ConfigMap/Secret name collision (it renames the ConfigMap with a `-config` suffix) and an `acceptAll*` domain needing a dedicated load balancer. Everything below changes or disappears **silently**, so diff the source against the output.
- **Scaling is pinned, not autoscaled.** A converted workload gets `minScale = maxScale =` the Deployment's `replicas` (or `1` if unset) with `capacityAI: false` — no headroom. An HPA, if present, supplies min/max and a CPU target. Raise `maxScale` above `minScale` for anything that should scale, keep customer-facing `minScale ≥ 2`, and consider Capacity AI (autoscaling-capacity skill).
- **Silently dropped from the pod spec** (the workload runs, but differently): `envFrom` (bulk ConfigMap/Secret env — re-add the keys as `env` or a mounted dictionary secret), `initContainers` (migrations/setup — run as a separate cron workload or an entrypoint step), `startupProbe` (only liveness/readiness carry over), container-level `securityContext` (only the pod-level `securityContext.fsGroup` carries over, as `filesystemGroupId`), and `hostPath` volumes. `emptyDir` becomes a `scratch://` volume.
- **Not converted at all** (no resource, no warning): NetworkPolicy, PodDisruptionBudget, RBAC, ResourceQuota/LimitRange, ServiceMonitor and other CRDs, and Namespaces — every namespace collapses into the one target GVC. Re-express network rules as the workload firewall (firewall-networking skill) and RBAC as policies (access-control skill).
- **Images stay literal, sizing is minimal.** `image: nginx:1.25` is kept verbatim, not rewritten to an internal `//image/` ref. `imagePullSecrets` (pod or ServiceAccount) carry over as `//secret/NAME`, but the secret must already exist for a private registry to pull. A container with no `resources` set defaults to a tiny `50m` CPU / `128Mi` memory — size it for production.
## Docker Compose (`cpln stack`)
`cpln stack deploy` (alias `up`) builds and deploys; `cpln stack manifest` previews the generated YAML without deploying; `cpln stack rm` (alias `down`) tears down. All take `--dir`/`--directory` and `--compose-file`. `--build` defaults `true` for `deploy` (local `docker build` for `linux/amd64`, then push as `<service>:1.0`) and `false` for `manifest`. Conversion rules:
- **Workload type:** `standard`, or `stateful` if the service attaches a named volume.
- **Volumes:** a named volume becomes a VolumeSet (stateful). A **file** bind mount becomes an opaque secret mounted at the target path. A **directory** bind mount is rejected with an error — split it into individual file bind mounts.
- **Ports:** `"PORT[:TARGET]/PROTO"` where `PROTO` is `http`, `http2`, `tcp`, or `grpc`; with no `/PROTO` suffix the port has no protocol set. Example: `"50051:50051/grpc"`.
- **Resources:** default `cpu: 42m`, `memory: 128Mi` (override via `deploy.resources.limits`). A GPU forces a minimum `cpu: 2000m`, `memory: 7168Mi`.
- **Firewall:** external inbound is opened (`0.0.0.0/0`) when the service has `ports` or `network_mode: host`; outbound is open unless `network_mode: none`.
- **Secrets/configs:** compose `secrets` and `configs` become opaque secrets with an auto-created identity and a `reveal` policy.
**`x-cpln` override block:** any top-level key under a service's `x-cpln` **replaces that entire `spec.<key>` section wholesale** (it does not deep-merge) — there is no allowlist, so any spec field works (`type`, `containers`, `defaultOptions`, `firewallConfig`, `identityLink`, …). Overriding `containers` means restating the full container spec.
```yaml
services:
api:
image: my-api:latest
x-cpln:
type: serverless # replaces the derived workload type
defaultOptions: # replaces the whole defaultOptions block
capacityAI: false
autoscaling: { minScale: 2, maxScale: 10 }
```
`cpln stack` does **not** rewrite service URLs in your code or config. Update them to the internal form `<workload>.<gvc>.cpln.local[:<port>]` (e.g. `http://redis:6379` becomes `http://redis.GVC.cpln.local:6379`).
## Bind-mounts: content vs config
The decisive question for any bind-mounted file: **does it change between environments, or is it identical in dev/staging/prod?**
| Type | Examples | Where it goes |
|---|---|---|
| **Application content** — versioned with the code, same everywhere | `index.html`, JS/CSS bundle, fonts, ML model weights | Bake into the image (`COPY` in the Dockerfile), or serve from a CDN |
| **Configuration** — env-specific values | `nginx.conf` with `proxy_pass`, app config with env URLs, `.env` | Opaque secret mounted as a file volume — the ConfigMap equivalent |
Baking config into the image couples that image to one environment: a hostname or feature-flag change then forces a rebuild. Mounting application content as secret volumes is the opposite mistake — it decouples content from its image version and makes rollbacks strange. A single container's migration is usually **mixed**, decided file by file. Rule of thumb: if changing the file between environments would not count as a code change, it is config and belongs in a secret volume.
Workloads mount secrets as read-only files via `cpln://secret/<name>` volumes (the `workload`/`stateful-storage` skills own the mechanism; `get_resource_schema` for `workload` gives the exact shape):
- **Opaque** (single config file): the path needs at least one subpath; the last segment becomes the file name and holds the `payload`.
- **Dictionary** (multi-key ConfigMap): mount the secret at a directory path and each key becomes a file.
- **Docker / GCP / Azure SDK**: mounted as a single file `___cpln___.secret` in the path directory.
So an nginx workload migrates *mixed*: `index.html` baked into a custom image, while `nginx.conf` mounts from a `cpln://secret/<name>` opaque secret at `/etc/nginx/conf.d/default.conf`.
## Helm (`cpln helm`)
`cpln helm install|upgrade|uninstall|list|template` manages releases of charts that render **only Control Plane kinds**. A rendered object carrying `apiVersion` or `metadata`, or an unknown kind, aborts with `ERROR: Some resources in the rendered template are not CPLN resources`. To migrate an existing Kubernetes Helm chart, render it first and pipe through the converter: `helm template R ./chart | cpln convert -f -`.
- `cpln.org` and `cpln.gvc` (plus `globals.cpln.*` / `global.cpln.*`) are injected as `--set` overrides — don't define a top-level `cpln` key in `values.yaml`, it gets clobbered.
- Release state is an opaque secret per revision; `cpln helm list` is org-scoped and takes no `--gvc`.
- GVC-scoped kinds (`workload`, `identity`, `volumeset`) need `--gvc` or a profile GVC; org-scoped kinds like `domain` do not.
Release-name rules, `--history-limit`, OCI charts, and `--wait` are general helm-release operations — the gitops-cicd and cpln skills own those.
## Exporting to Terraform / IaC
When the target is Infrastructure-as-Code rather than live resources, turn the converted Control Plane YAML into HCL with `convert_to_terraform` (dry-run validated against the API first, so the HCL always matches a schema-valid resource), or capture already-created resources with `export_terraform`. `list_terraform_kinds` and `export_terraform_batch` are in the `full` profile. The `iac-terraform-pulumi` skill owns the full Terraform/Pulumi story, including `terraform import`.
## Verify
- After `cpln convert`: confirm each workload's derived type, scaling (`maxScale` raised where needed), port protocols (gRPC/HTTP2), ingress-to-domain routes, and that any `{{GVC}}` placeholder is replaced.
- After create/apply: `cpln apply -f cpln.yaml --ready`, or poll `list_deployments` until each workload reports ready. Pair every mutation with a read.
## Troubleshooting
| Symptom | Cause and fix |
|---|---|
| `cpln helm`: "…not CPLN resources" | Chart renders Kubernetes objects (`apiVersion`/`metadata`). Render then convert: `helm template \| cpln convert`. |
| Env vars missing, or a setup step never ran | `envFrom` and `initContainers` are dropped silently — re-add env keys as `env`/a dictionary secret, and run init logic as a cron workload or entrypoint step. |
| Workload won't scale under load | `minScale = maxScale` from the source replicas — raise `maxScale` (and enable Capacity AI / a metric). |
| Compose: "Directory bind mount found" | Directory bind mounts are rejected — mount individual files (each becomes a secret). |
| App can't reach another service | The converters don't rewrite URLs — point them at `<workload>.<gvc>.cpln.local[:port]`. |
| Private image won't pull | The image string is kept literal; create the pull secret it references and link it (image skill). |
| Deployment stuck after converting | The workload references a secret without an identity/policy; the converter adds `identity-<wl>`/`policy-<wl>` — if you re-authored, wire `reveal` yourself (access-control skill). |
## Quick reference
### MCP tools
- `create_workload` / `create_gvc` / `create_identity` / `create_volumeset` — author converted resources with production-grade defaults (secrets are created by the user — draft each manifest for them to fill and apply, `setup-secret` skill; verify with `get_resource` before referencing)
- `get_resource_schema` — exact shape before hand-editing or re-authoring a converted manifest
- `list_deployments` — poll converted workloads to ready
- `convert_to_terraform` / `export_terraform` — converted YAML or live resources to HCL (`iac-terraform-pulumi` skill)
The converters themselves (`cpln convert`, `cpln stack`, `cpln helm`, `cpln apply --k8s`) are CLI-only. In CI/CD, `CPLN_TOKEN` + `cpln apply -f` applies the converted manifest headlessly.
### Related skills
| Skill | Use for |
|---|---|
| workload | the spec the converter emits; deploy/diagnose flow, injected `CPLN_*` vars |
| cpln | the CLI that runs every converter; `apply` ordering, `exec`/`logs` |
| autoscaling-capacity | giving converted workloads scaling headroom and Capacity AI |
| stateful-storage | volumeset shape for converted PVCs and compose named volumes |
| iac-terraform-pulumi | turning converted YAML into Terraform or Pulumi |
| template-catalog | deploy a database from a template instead of converting one |
## Documentation
- [cpln convert](https://docs.controlplane.com/guides/cli/cpln-convert.md)
- [Compose Deploy](https://docs.controlplane.com/guides/compose-deploy.md)
- [cpln helm](https://docs.controlplane.com/guides/cpln-helm.md)
- [cpln apply](https://docs.controlplane.com/guides/cpln-apply.md)
- [Workload Volumes](https://docs.controlplane.com/reference/workload/volumes.md)
mk8s-byok16.2 KB
--- name: mk8s-byok description: "Runs workloads on your own hardware and provisions managed Kubernetes (mk8s) clusters on Control Plane. Use when the user asks about bare metal, on-prem, data centers, their own servers, mk8s, BYOK, or node pools." --- # Managed Kubernetes (mk8s) & BYOK > **Tool availability:** the `create_mk8s_*` / `update_mk8s_*` tools live in the `mk8s` toolset profile (`?toolsets=mk8s`; `full` includes it). If an mk8s tool is not advertised, tell the user to reconnect with `?toolsets=mk8s`. Provider credential secrets (opaque token, gcp, keypair) must already exist — created by the user; offer to draft the manifest for them to fill and apply (`setup-secret` skill). Reads work on every profile via `list_resources` / `get_resource` (kind `mk8s` or `location`); `delete_resource` is on every profile except `readonly`. Control Plane has three separate "Kubernetes" stories people routinely conflate — get the right one first: - **mk8s** — Control Plane *provisions and manages* a real, conformant Kubernetes cluster on your cloud account (12 providers). You get a kubeconfig and run normal Kubernetes. It is a **standalone cluster** — to schedule Control Plane (GVC) workloads onto it, add the `byok` add-on (below). Resource kind `mk8s`. - **BYOK location** — you already have a self-managed cluster; you *register it* as a Control Plane location so Control Plane workloads (GVC workloads) schedule onto it. Resource kind `location`, provider `byok`. - **mk8s BYOK add-on** — `addOns.byok` makes an mk8s cluster register *itself* as a Control Plane location: it links a `location` you create and installs the agent automatically (no manual `cpln location install`). This is how Control Plane workloads run on an mk8s cluster. (If the user instead wants to manage Control Plane resources *from* `kubectl`, that is the **k8s-operator** skill, not this one.) The dominant failures: reaching for a nonexistent `cpln mk8s create`; skipping the per-provider credential secret; and leaving the cluster's API-server firewall wide open. ## "Can I run this on my own servers?" — yes Bare metal in a data center or colo, on-prem VMs (VMware/vSphere), a Dell or Supermicro rack — any Linux server becomes a Control Plane **location** the user deploys to exactly like `aws-us-east-1`. "BYOC" means the same thing with the cluster in their own cloud account. Never answer this question with cloud regions only. Two routes, picked by what they already have: | They have | Route | |:---|:---| | A Kubernetes cluster already (EKS, GKE, AKS, k3s, self-managed) | Register it as a **BYOK location** (below) | | Servers but no cluster | Build one with the **`generic`** provider (`create_mk8s_generic`, then `cpln mk8s join` per node), then register it | Both end at a location. From there it is ordinary workload work: add the location to a GVC (`update_gvc` with `addLocations`), deploy, verify. **The BYOK prerequisites are the binding constraint, not the mk8s ones** — a generic node needs only 1 CPU / 512 MB, but a cluster serving as a *location* needs ≥ 2 nodes, ≥ 2 CPU and 8 GB each, and a working LoadBalancer controller. Say so before the user buys hardware. ## Providers & credential secrets Exactly one provider per cluster (the schema enforces XOR). For every provider except AWS and Generic, **create the credential secret first**, then reference it in the create call. | Provider | Create first (credential) | Node-sizing field | Location/region | |:---|:---|:---|:---| | `aws` | `deployRoleArn` — an assumed IAM role, **no secret** (also needs `vpcId`) | `instanceTypes[]` | `region` | | `azure` | **opaque** secret (`sdkSecretLink`) — service-principal creds | `size` | `location` | | `gcp` | **gcp** secret (`saKeyLink`) — SA JSON key | `machineType` | `region` | | `digitalocean` | **opaque** secret (`tokenSecretLink`) | `dropletSize` | `region` | | `hetzner` | **opaque** secret (`tokenSecretLink`) | `serverType` | `region` | | `linode` | **opaque** secret (`tokenSecretLink`) | `serverType` | `region` | | `oblivus` | **opaque** secret (`tokenSecretLink`) | `flavor` (GPU enum) | `datacenter` | | `lambdalabs` | **opaque** secret (`tokenSecretLink`) | `instanceType` (GPU enum) | `region` | | `paperspace` | **opaque** secret (`tokenSecretLink`) | `machineType` (GPU enum) | `region` | | `triton` | **keypair** secret (`connection.privateKeySecretLink`) | `packageId` | `location` | | `generic` | none — you join your own nodes | n/a (external nodes) | `location` | **Azure uses an `opaque` secret, not an `azure-sdk` secret** — the create tool rejects the typed one. Triton/Generic/Ephemeral `location` is a *Control Plane* location (e.g. `aws-us-east-2`), not a cloud region. A 12th provider, `ephemeral`, exists in the schema but has **no create tool or CLI create** — ignore it for real clusters. Other required fields (network/VPC, image, SSH keys, region enum) vary per provider; `mcp__cpln__get_resource_schema` (kind `mk8s`) and the create tool's own validation give the exact required set. Generic nodes you supply each need Linux kernel ≥ 5.4, ≥ 1 CPU / 512 MB, mutual connectivity, and SSH access. ## The cluster spec essentials - **`version`** — required, no default, a **closed enum** of specific patch versions (currently `1.26.0` through `1.35.3`). The set drifts as versions are added/retired; pull the live list from `get_resource_schema`, or just submit and let the typed tool's validation error name the valid values. Updating it is **upgrade-only** — mk8s rejects downgrades. - **`nodePools`** (per provider) — give at least one to get worker capacity. Common fields are `name`/`labels`/`taints`; the rest are provider-specific. Each pool's **`minSize`/`maxSize` drive the cluster autoscaler** (set `maxSize > minSize` to allow scale-up). - **`autoscaler`** (per provider) — defaults: `expander: [most-pods]`, `unneededTime: 10m`, `unreadyTime: 20m`, `utilizationThreshold: 0.7`. Usually leave it. - **`networking`** (per provider) — `serviceNetwork` (default `10.43.0.0/16`) and `podNetwork` (default `10.42.0.0/16`) must differ and are **unmodifiable after creation**; pick non-overlapping CIDRs up front. AWS and GCP additionally accept `podNetwork: vpc`. - **`firewall`** — the API-server allow-list. It **defaults to `0.0.0.0/0` (fully open)**, the opposite of the deny-by-default workload firewall; the create tool warns when you omit it. Restrict it to your admin CIDRs. Each rule is `sourceCIDR` (required) + `description`. ## Add-ons `spec.addOns` is a map of toggles and small config objects. The schema is **provider-agnostic — it does not block any add-on on any provider**; an add-on only *functions* where the provider supports it (column below is the documented applicability). | Add-on key | What it does | Providers | |:---|:---|:---| | `dashboard` | Kubernetes Dashboard UI | all | | `headlamp` | Headlamp web UI | all | | `metrics` | Prometheus metrics to Control Plane | all | | `logs` | pod logs + audit records to Control Plane | aws, hetzner, generic | | `localPathStorage` | local-path PVC provisioner | all | | `registryMirror` | P2P image-layer cache across nodes | all | | `sysbox` | run Docker/Kubernetes inside pods (stronger isolation) | all | | `nvidia` | NVIDIA GPU operator (`taintGPUNodes`) | GPU nodes | | `kubevirt` | run VMs on the cluster (needs `nodeLocalDns` + `cpln.io/nodeType=vm` nodes) | nodes with HW virtualization | | `nodeLocalDns` | per-node CoreDNS cache | all | | `awsWorkloadIdentity` | pods assume AWS IAM roles | aws, hetzner, generic | | `awsECR` | pull from private ECR | aws, hetzner, generic | | `awsEFS` | EFS volumes (`roleArn` required) | aws, hetzner, generic | | `awsELB` | AWS Load Balancer Controller (NLB/ALB) | aws | | `azureWorkloadIdentity` | pods get Azure AD identities | all | | `azureACR` | pull from private ACR (`clientId` required) | all | | `byok` | register this cluster as a BYOK location (`addOns.byok.location` is a `//location/NAME` link) | all | ## Create & update (MCP) Pick the provider's tool — `mcp__cpln__create_mk8s_<provider>` for any provider in the table above (e.g. `create_mk8s_aws`, `create_mk8s_generic`). Supply `version` and at least one node pool. Read with `mcp__cpln__get_resource` / `list_resources` (kind `mk8s`); delete with `mcp__cpln__delete_resource` (destructive — confirm the blast radius). `mcp__cpln__update_mk8s_<provider>` is a **merge-patch**: send only what changes. Provided **arrays** (`nodePools`, `firewall`) **replace the previous values wholesale** — include the full set you want to keep. `addOns` merge-patches per key, so the update tool **cannot disable a toggle add-on by passing `false`** — remove one with `cpln apply` (set the key to `null`). `region`/`networking` and the provider/name are unmodifiable after creation. Fallback (MCP unavailable, or CI/CD): author a YAML manifest from `get_resource_schema` (kind `mk8s`) and `cpln apply --file mk8s.yaml`. ## CLI (mk8s) There is **no `cpln mk8s create`** — create via the MCP tools above or `cpln apply`. The CLI owns the operations with no MCP equivalent: | Command | Purpose | |:---|:---| | `cpln mk8s get [ref...]` / `query` | read clusters | | `cpln mk8s health <ref>` | readiness status of the cluster | | `cpln mk8s kubeconfig <ref>` | generate a kubeconfig | | `cpln mk8s dashboard <ref>` | open the Kubernetes dashboard | | `cpln mk8s join <ref>` | join your own nodes (generic clusters and Hetzner dedicated-server pools) | | `cpln mk8s eventlog <ref>` | cluster event log (alias `log`) | | `cpln mk8s clone <ref>` | duplicate the spec (alias `copy`) | | `cpln mk8s edit / patch / update / delete` | edit YAML / patch metadata / `--set` / remove | `cpln mk8s update --set` accepts only `description`, `tags.<key>`, and `spec.version`. For provider, node-pool, or add-on changes use `update_mk8s_<provider>`, `cpln mk8s edit`, or `cpln apply`. ## BYOK location (register an existing cluster) For a cluster you already run yourself. **All location create/install/uninstall steps are CLI-only — there is no MCP tool.** 1. `cpln location create --name CLUSTER` — create the BYOK location entry. 2. `cpln location install CLUSTER` — prints instructions for obtaining the install script (a signed `kubectl apply` command). The manifests carry sensitive tokens and are valid for **about 5 minutes** — run it promptly or regenerate. 3. Apply it against the target cluster's kubectl context. 4. Wait for the **`cpln-byok-agent`** deployment in the **`kube-system`** namespace to become ready: `kubectl get pod -l app=cpln-byok-agent -n kube-system`. 5. Add the location to a GVC, then deploy workloads onto it. Remove with `cpln location uninstall CLUSTER` and run the printed command on the cluster. **Prerequisites:** a cluster within the three most recent Kubernetes minor releases; ≥ 2 nodes (3+ recommended); ≥ 2 CPU and 8 GB RAM per node (4 / 16 recommended); architecture `amd64` or `arm64`; at least one nodegroup labeled **`cpln.io/nodeType=core`**; full node-to-node connectivity and egress; a working LoadBalancer controller (a `Service` of type LoadBalancer must obtain an IP); and **no pre-installed service mesh** — Control Plane installs its own Istio-based mesh. **Provider notes.** GKE: first give the `kube-dns` Service IP (`kubectl get svc -n kube-system kube-dns`) to support; then, *after* Control Plane config is applied, scale `kube-dns` and `kube-dns-autoscaler` (in `kube-system`) to 0 replicas. EKS: enable the Amazon VPC CNI, `kube-proxy`, CoreDNS, and Amazon EBS CSI Driver add-ons. On-prem/airgapped: contact support. Once the location exists, prefer MCP for the GVC and workload work: `mcp__cpln__update_gvc` with `addLocations: [CLUSTER]`, then deploy and poll `mcp__cpln__list_deployments`. CLI fallback: `cpln gvc add-location GVC --location CLUSTER`. ## Verify - **mk8s:** poll `mcp__cpln__get_resource` (kind `mk8s`) until `status.serverUrl` is set, or `cpln mk8s health CLUSTER`; then `cpln mk8s kubeconfig CLUSTER` and `kubectl get nodes` to confirm worker capacity. - **BYOK:** the agent pod is ready (`kubectl get pod -l app=cpln-byok-agent -n kube-system`), the location reports enabled, and a test workload targeting the location reaches ready in `list_deployments`. ## Troubleshooting | Symptom | Cause and fix | |:---|:---| | Version create/update rejected | `version` must be in the closed enum, and updates are **upgrade-only** (no downgrade); read the valid set from `get_resource_schema` or the error. | | Create rejected: secret type | Azure needs an **opaque** secret (not `azure-sdk`); GCP needs a **gcp** secret; Triton a **keypair** secret; token providers an **opaque** secret. Create it first. | | Cluster reachable from anywhere | The API-server `firewall` defaulted to `0.0.0.0/0` — set it to your admin CIDRs. | | Update wiped node pools / a rule | `nodePools` and `firewall` arrays replace wholesale on update — resend the full set. | | Can't disable an add-on via update | The update tool ignores `false` toggles — remove the key with `cpln apply` (`null`). | | Generic cluster has no workers | Generic nodes are external — run `cpln mk8s join` on each node. | | BYOK agent never readies | Missing `cpln.io/nodeType=core` nodegroup, no working LoadBalancer, a pre-existing service mesh, or the install command expired (~5 min) — regenerate with `cpln location install`. | | BYOK on GKE: DNS conflicts | Scale `kube-dns` and `kube-dns-autoscaler` to 0 and hand the `kube-dns` IP to support. | ## When this is the wrong tool - **One server.** A single box can be a generic *worker node*, but a BYOK *location* expects ≥ 2 nodes and a LoadBalancer controller. Do not promise a one-machine deployment target — say what the minimum actually is. - **They only need to *reach* something on-prem.** A workload running on Control Plane that must talk to a data-center database or internal API needs a wormhole agent (`setup-agent`) or PrivateLink/PSC (`native-networking`) — not a location on their hardware. "On-prem" in a question about *connectivity* is a different skill; this skill is for on-prem *compute*. - **Air-gapped or no egress.** Nodes require egress access; air-gapped installs are a support conversation, not a self-service path. - **Managing Control Plane from `kubectl`.** That is `k8s-operator`, the opposite direction. - **They just want a container running somewhere.** If the user never asked for their own hardware, a cloud location is simpler — do not route them through a cluster build. ## Quick reference ### MCP tools - `mcp__cpln__create_mk8s_<provider>` / `update_mk8s_<provider>` — create/merge-patch a cluster (mk8s profile) - `mcp__cpln__get_resource` (kind `secret`) — verify the provider credential secret exists (created by the user) - `mcp__cpln__get_resource` / `list_resources` / `delete_resource` (kind `mk8s` or `location`) — read/delete on any profile - `mcp__cpln__get_resource_schema` (kind `mk8s`) — exact shape and the live `version` set before authoring YAML - `mcp__cpln__add_gvc_locations` / `list_deployments` — attach a BYOK location to a GVC and verify workloads BYOK *location* create/install/uninstall and `cpln mk8s kubeconfig|join|dashboard|health` are **CLI-only**. In CI/CD, `CPLN_TOKEN` + `cpln apply -f mk8s.yaml` provisions a cluster headlessly. ### Related skills | Skill | Use for | |:---|:---| | workload | deploying workloads onto the cluster once its location is in a GVC | | cpln | the CLI behind `mk8s` and `location` (kubeconfig, join, install) and `cpln apply` | | stateful-storage | volumesets and the BYOK volumeset storage-class settings | | access-control | policies and grantable permissions on cluster/location objects | | image | pull secrets behind the ECR/ACR add-ons | | k8s-operator | the opposite direction — managing Control Plane resources from `kubectl` | ## Documentation - [Deploy a Workload to Your Own Hardware](https://docs.controlplane.com/guides/deploy-to-your-own-hardware.md) - [mk8s Overview](https://docs.controlplane.com/mk8s/overview.md) - [mk8s on AWS](https://docs.controlplane.com/mk8s/aws.md) · [GCP](https://docs.controlplane.com/mk8s/gcp.md) · [Hetzner](https://docs.controlplane.com/mk8s/hetzner.md) · [Triton](https://docs.controlplane.com/mk8s/triton.md) · [Generic](https://docs.controlplane.com/mk8s/generic.md) - [BYOK Overview](https://docs.controlplane.com/byok/overview.md) - [CLI mk8s Commands](https://docs.controlplane.com/cli-reference/commands/mk8s.md)
native-networking13 KB
---
name: native-networking
description: "Connects Control Plane workloads to private VPCs, on-prem networks, and cross-cloud resources. Use when the user asks about AWS PrivateLink, GCP Private Service Connect, wormhole agents, or reaching a private network."
---
# Native Networking & Agent Connectivity
A Control Plane workload reaches a private or cross-cloud endpoint through an **identity** (gvc-scoped) carrying one of two resource arrays. Attach that identity to the workload (`spec.identityLink`) — without the attachment, nothing routes. Both paths are wired **independently of the workload's external egress firewall**: you do *not* open an `outboundAllow*` rule to reach them. The two options:
- **Native networking** (`nativeNetworkResources`) — cloud-native private connectivity over **AWS PrivateLink** or **GCP Private Service Connect**. No agent, lowest latency, no public-internet traversal. The catch: the consumer-side endpoint is created by **Control Plane support**, not self-service.
- **Agent / wormhole** (`networkResources`) — a lightweight VM or container you run inside the target network that tunnels TCP traffic. Self-service, works for **any** network (VPC, on-prem, cross-cloud, Azure, a laptop), but throughput depends on the agent instance size.
> **Scope:** this skill is the reference for the comparison, producer-side setup, the identity schema, agent sizing, and permissions. For the agent **deployment walkthrough** (create, generate K8s/Docker/VM artifacts, wire up the identity, verify the tunnel), delegate to the **setup-agent** skill.
## Choosing an option
| Target | Option | Agent? | Consumer side set up by |
|:---|:---|:---|:---|
| AWS service (RDS, etc.) | AWS PrivateLink (native) | No | Support, then **you accept** the endpoint in the AWS console |
| GCP service (Cloud SQL, etc.) | GCP Private Service Connect (native) | No | Support (Cloud SQL needs **no** manual acceptance) |
| On-prem / data center | Agent | Yes | Self-service |
| Cross-cloud / multi-VPC | Agent | Yes | Self-service |
| Azure VNet, or a developer laptop | Agent | Yes | Self-service |
**This skill is about *reaching* a private network, not running in one.** If the user wants the workload itself to *run* on their own hardware — bare metal, an on-prem VM, a data-center server — that is a BYOK location, not an agent: see `mk8s-byok`.
## Calling a resource from a workload
Once the identity is attached (`spec.identityLink`), the workload reaches either kind of resource like an ordinary host — no SDK, env var, or code change:
- **Connect to the resource's `name`** (or its `FQDN`) on one of the configured **`ports`** — e.g. a Postgres client points at `aws-rds:5432` (native) or `on-prem-db:5432` (agent).
- Control Plane injects a hosts entry so that name resolves and routes to the real endpoint: for **native**, straight to the PrivateLink/PSC private IP; for an **agent**, through the tunnel to the upstream `IPs`/`FQDN` on the private side.
- **Use the `FQDN`, not the `name`, when the target serves TLS** — the certificate is issued for the FQDN, so the short `name` fails certificate validation.
- Only the ports you list are wired to the resource — a port you did not configure is not opened.
## Native networking (PrivateLink / PSC)
Traffic flows from the workload, through Control Plane infrastructure, to your cloud's private endpoint — never the public internet. Setup:
1. **Provision the producer side.** Use the reference Terraform, or wire up an existing resource:
- AWS RDS + PrivateLink: `github.com/controlplane-com/cpln-rds-producer` (new-infra mode also builds the VPC/RDS/Secrets Manager; existing-infra mode adds only RDS Proxy + NLB + Lambda + the endpoint service). Output: the **endpoint service name**.
- GCP Cloud SQL + PSC: `github.com/controlplane-com/gcp-psc-producer-automation`. Output: the **service attachment**. For an *existing* Cloud SQL instance, enable PSC via gcloud — it is **not available in the GCP console**:
```bash
gcloud sql instances patch INSTANCE --enable-private-service-connect --allowed-psc-projects=cpln-prod01
```
The allowed consumer project must be **`cpln-prod01`**. Cloud SQL must use a private IP only.
2. **Hand the service name (AWS) / service attachment (GCP) plus the region to `support@controlplane.com` (or ping support on Slack).** They create the consumer-side endpoint and associate it with your org.
3. **AWS only:** accept the connection in the AWS console (VPC, Endpoint Services, Pending endpoint connections, Accept). Cloud SQL connections are accepted automatically.
4. **Add a `nativeNetworkResources` entry** to the identity (tools and schema below), then attach the identity to the workload.
> **The identity entry is inert until support has wired the consumer side** (and, for AWS, you have accepted the endpoint connection). Until then — or if the `endpointServiceName` is mistyped — the platform **silently skips** it (no error, no connection). So add the entry *last*, not first.
```yaml
nativeNetworkResources:
- name: "aws-rds" # a label; must be unique and must NOT equal the FQDN
FQDN: "rds-proxy.us-west-2.amazonaws.com"
ports: [5432]
awsPrivateLink:
endpointServiceName: "com.amazonaws.vpce.us-west-2.vpce-svc-12345abcdef"
- name: "gcp-sql"
FQDN: "my-sql.us-central1.gcp.internal"
ports: [5432]
gcpServiceConnect:
targetService: "projects/PROJECT/regions/us-central1/serviceAttachments/NAME"
```
## Agents (wormholes)
An agent runs inside the target network and opens a persistent **outbound** connection to Control Plane; workload requests are tunneled through it (workload, Control Plane, agent, private endpoint). Use it for on-prem, cross-cloud, Azure, or local development — anywhere PrivateLink/PSC does not reach.
The deployment flow (create the agent, deploy it, attach `networkResources`, verify) is owned by the **setup-agent** skill. The pieces that belong here regardless of how it is deployed:
```yaml
networkResources:
- name: "on-prem-db"
agentLink: "//agent/dc-agent" # or /org/ORG/agent/dc-agent
IPs: ["10.0.1.50"] # OR FQDN — exactly one
ports: [5432, 3306]
```
**High availability** (`reference/agent.mdx`): run agents in a fixed-size instance group (autoscaling group on AWS, VMSS on Azure) sized **min 2, max = number of availability zones**. The agent is **not CPU-intensive — do not autoscale on CPU.** Agents run **active-active**: every instance registers and serves traffic at once (load-balanced); if one misses heartbeats it is dropped and the rest keep serving while the group replaces it.
**Bi-directional:** the agent also exposes a proxy on port **3128** so systems inside the private network can call Control Plane workloads without opening external firewall access — enable with `cpln agent up --exposeProxy` and grant it on the workload's **Internal** firewall (Add Agent — see `firewall-networking`).
### Agent sizing
Benchmarked with qperf (30s) from a workload in the server VM's region. Plan against **baseline** bandwidth, not burst.
**AWS** — server `c5.2xlarge` in `aws-us-west-2`:
| Agent instance | Bandwidth (MB/s) | Latency (us) | Baseline (Gbps) |
|:---|:---|:---|:---|
| No agent | 307.6 | 585.6 | n/a |
| t2.micro | 21.23 | 1301 | 0.064 |
| t3.small | 143.9 | 1107 | 0.128 |
| c5.large | 341.1 | 629.6 | 0.75 |
| c4.xlarge | 70.25 | 680.8 | 5.0 |
Find any instance's baseline with `aws ec2 describe-instance-types --query "InstanceTypes[].[InstanceType,NetworkInfo.NetworkCards[0].BaselineBandwidthInGbps]"`.
**GCP** — server `e2-standard-8` in `gcp-us-east1`:
| Agent machine | Bandwidth (MB/s) | Latency (us) |
|:---|:---|:---|
| No agent | 313.4 | 251.2 |
| n2-standard-2 | 250.3 | 407.7 |
| n2-standard-8 | 223.3 | 350.7 |
| n2-standard-4 | 217.5 | 354.1 |
| n1-standard-1 | 199.9 | 409.3 |
## Identity network-resource schema
Both arrays live on the identity object. Constraints are enforced by the platform (Joi) — the typed tools mirror them.
**`networkResources`** (agent-based):
| Field | Required | Notes |
|:---|:---|:---|
| `name` | Yes | label or domain; the dialable hostname |
| `agentLink` | No | `//agent/NAME` or `/org/ORG/agent/NAME` |
| `IPs` | one of | 1-5 IPv4 — **xor with `FQDN`** |
| `FQDN` | one of | one domain — **xor with `IPs`** |
| `resolverIP` | No | IPv4 of the DNS resolver the agent uses to resolve the `FQDN` inside the private network |
| `ports` | Yes | 1-10 ports, each 0-65535 |
**`nativeNetworkResources`** (PrivateLink / PSC):
| Field | Required | Notes |
|:---|:---|:---|
| `name` | Yes | label; must not equal the FQDN |
| `FQDN` | No | use it for TLS targets |
| `ports` | Yes | 1-10 ports, each 0-65535 |
| `awsPrivateLink.endpointServiceName` | one of | **xor with `gcpServiceConnect`** |
| `gcpServiceConnect.targetService` | one of | `projects/…/regions/…/serviceAttachments/…` — **xor with `awsPrivateLink`** |
Global rules (Joi-enforced): each array holds **max 50** entries, and **`name` and `FQDN` must be unique across both arrays combined**. Operationally (not a schema rule), two native resources that share a port need separate PrivateLink/PSC endpoints.
## Configuring & verifying
- Attach resources with `add_identity_native_network_resource` (PrivateLink/PSC) or `add_identity_network_resource` (agent-based) — each takes `org`, `gvc`, `identity`, and one `resource`. Create the identity first (`create_identity`, see `access-control`) if it does not exist.
- `remove_identity_network_resource` removes from **either** array by name (destructive — present the impact and get the user's explicit approval before calling). There is no separate native-remove tool.
- **Verify:** `list_identity_network_resources` lists both arrays; `get_agent_info` shows whether an agent is active plus its `peerCount` / `serviceCount`; then confirm the workload actually connects to the endpoint.
## Agent permissions
| Permission | Grants | Implies |
|:---|:---|:---|
| `view` | read-only | |
| `use` | reference the agent in an identity | view |
| `edit` | modify the agent | view |
| `create` | create agents | |
| `delete` | delete agents | |
| `manage` | full access | create, delete, edit, use, view |
## Troubleshooting
| Symptom | Cause and fix |
|:---|:---|
| Native resource never connects (no error) | Support hasn't created the endpoint / allow-listed the service, (AWS) the connection wasn't accepted, or `endpointServiceName` is mistyped — the platform silently skips an unmatched native resource. |
| TLS / certificate error to a native resource | Connect via the `FQDN`, not the `name` — the cert is issued for the FQDN. |
| "Provide exactly one of…" on the entry | `IPs` xor `FQDN` (agent), or `awsPrivateLink` xor `gcpServiceConnect` (native) — supply exactly one. |
| Duplicate name/FQDN rejected | Names and FQDNs must be unique across **both** arrays; the `name` must not equal any FQDN. |
| Agent shows inactive (`get_agent_info`) | No recent heartbeat — the deployed agent is down or cannot reach Control Plane; check `get_agent_eventlog`. |
| Workload cannot reach the agent's network | The identity is not attached to the workload (`spec.identityLink`), or the resource ports are wrong. |
| Agent will not delete | It is still referenced by an identity — remove the `networkResource` (or detach the identity) first. |
| Bootstrap config lost | It is shown only at `create_agent` time and is immutable — delete and recreate the agent. |
## Quick reference
### MCP tools
- `create_agent` / `update_agent` — create (returns the one-time bootstrap config) / patch description & tags (full profile)
- `get_agent_info` / `get_agent_eventlog` — live status and event log (full profile)
- `add_identity_native_network_resource` / `add_identity_network_resource` — attach a PrivateLink/PSC or agent resource (full profile)
- `remove_identity_network_resource` / `list_identity_network_resources` — remove from either array / list both (full profile)
- `get_resource` / `list_resources` / `delete_resource` (kind `agent` or `identity`) — read and delete on any profile
CLI fallback (MCP unavailable, or CI/CD with `CPLN_TOKEN`): `cpln agent create|manifest|up|info|eventlog`. Network resources on an identity are **not** settable via `cpln identity create`/`update` (description and tags only) — edit the identity YAML with `cpln identity edit REF` or `cpln apply -f identity.yaml`.
### Related skills
| Skill | Use for |
|:---|:---|
| workload | attaching the identity to the workload (`spec.identityLink`) that needs the connectivity |
| access-control | creating the identity, and policies/permissions on agents |
| firewall-networking | the Internal firewall for the bi-directional proxy, and service-to-service rules |
| cpln | the `cpln agent` CLI and `cpln apply` |
### Documentation
- [Native Networking Setup](https://docs.controlplane.com/guides/native-networking/native-networking-setup.md)
- [Agent Reference](https://docs.controlplane.com/reference/agent.md) · [Agent Setup Guide](https://docs.controlplane.com/guides/agent.md)
- [Identity Reference](https://docs.controlplane.com/reference/identity.md)
query-spec5.87 KB
---
name: query-spec
description: "Filters, selects, and sorts Control Plane resources with the query spec language. Use when the user asks about targetQuery, memberQuery, cpln query commands, tag-based selection, property filtering, or dynamic location selection."
---
# Query Spec — Filtering & Selecting Resources
Control Plane has one query language used in two ways: **ad-hoc filtering** of a resource list (CLI / API), and **dynamic targeting** embedded inside three resource fields.
| Where | Field / command | Purpose |
|:---|:---|:---|
| Policy | `targetQuery` | Target resources by tag/property instead of listing `targetLinks` (see **access-control**) |
| Group | `memberQuery` | Assign members dynamically — **users only** |
| GVC | `spec.staticPlacement.locationQuery` | Select locations dynamically instead of listing `locationLinks` |
| CLI | `cpln KIND query` | Ad-hoc filtering — every resource kind |
| API | `POST /org/ORG/KIND/-query` | Ad-hoc filtering — every resource kind |
**Not an MCP list parameter.** `list_resources` has no filter argument (list a kind, then filter the table yourself); `query_audit_events` filters by kind/name/subject/context/time; `query_metrics` takes PromQL. The query spec appears only inside the three resource fields above, set when you create or update that resource.
## Structure
```yaml
kind: workload # resource kind being selected
fetch: items # "items" (objects, default) or "links" (references)
spec:
match: all # "all" (default), "any", or "none"
terms:
- op: "="
tag: environment
value: production
- op: exists
tag: monitored
sort:
by: name
order: asc
```
## Terms
Each term targets exactly **one** of three fields (mutually exclusive — the schema rejects a term that sets more than one):
| Field | Targets | Example |
|:---|:---|:---|
| `tag` | Resource tags (key/value labels) | `tag: environment`, `value: production` |
| `property` | Built-in properties (`name`, `description`, `status.phase`, …) | `property: name`, `value: my-app` |
| `rel` | Relationships to other resources | `rel: gvc`, `value: my-gvc` |
### Operators
| Operator | Needs `value` | Meaning |
|:---|:---:|:---|
| `=` | yes | Equal (default when `op` is omitted) |
| `!=` | yes | Not equal |
| `>` `>=` `<` `<=` | yes | Numeric / date comparison |
| `~` | yes | Pattern match (schema op name `match`) |
| `=~` | yes | Regex match (schema op name `regex`) |
| `contains` | yes | Substring match |
| `exists` | no | Tag/property is present (any value) |
| `!exists` | no | Tag/property is absent |
`value` accepts a string, number, boolean, or ISO date. **Boolean values are auto-converted to strings on `tag` terms only** — store `monitored=true` and you must query `value: "true"`, not `value: true`, or it silently matches nothing.
## Match modes
| Mode | Behavior |
|:---|:---|
| `all` | Every term must match (default) |
| `any` | At least one term matches |
| `none` | No term may match |
## Sorting
```yaml
sort:
by: name # required
order: asc # "asc" (default) or "desc"
```
Sort is **API- and manifest-only** — the `cpln KIND query` CLI has no sort flag, so a sort directive passed there is ignored.
Common fields (most kinds): `id`, `name`, `version`, `description`, `created`, `lastModified`. Kind-specific: `location` adds `origin`/`provider`/`region`; `cloudaccount` adds `provider`; `user` adds `idp`/`email`; `policy` and `group` add `origin`.
## CLI
**Ad-hoc filtering** — every kind supports `query`:
```bash
cpln workload query --tag environment=production
cpln workload query --match all --tag environment=production --tag region=europe
cpln workload query --rel gvc=my-gvc
cpln policy query --prop name=my-policy
cpln workload query --tag monitored # existence (no value)
cpln workload query --match any --rel gvc=one --rel gvc=two
```
| Flag | Alias | Notes |
|:---|:---|:---|
| `--match` | | `all` / `any` / `none` (default `all`); single value |
| `--tag` | | `KEY=VALUE`, or `KEY` for existence; repeatable |
| `--property` | `--prop` | `KEY=VALUE`; repeatable |
| `--rel` | | `KEY=VALUE`; repeatable |
Results cap at 50 by default — raise with `--max 0` for all records.
**Authoring dynamic targeting** — `gvc`, `policy`, and `group` create/update commands embed a query via `--query-match`, `--query-tag`, `--query-property`, `--query-rel` (group also `--query-kind user`):
```bash
cpln policy create --name img-policy --query-kind image --query-property repository=my-app ...
```
## API
```
POST https://api.cpln.io/org/ORG/workload/-query
```
```json
{ "spec": { "match": "all",
"terms": [
{ "op": "=", "tag": "region", "value": "emea" },
{ "rel": "gvc", "op": "=", "value": "mygvc" }
],
"sort": { "by": "name", "order": "asc" } } }
```
## Dynamic targeting examples
**Policy `targetQuery`** — applies to matching resources, including ones created later:
```yaml
kind: policy
targetKind: image
targetQuery:
spec:
terms:
- { property: repository, value: my-app }
bindings:
- permissions: [pull, view]
principalLinks: [//group/developers]
```
**Group `memberQuery`** — dynamic membership by user tag (users only; service accounts must be added via `memberLinks`):
```yaml
kind: group
memberQuery:
kind: user
spec:
terms:
- { tag: "firebase/sign_in_provider", value: "microsoft.com" }
```
## Defaults & gotchas
- Omitted `op` defaults to `=`; omitted `match` defaults to `all`; omitted `fetch` defaults to `items`; omitted sort `order` defaults to `asc`.
- Boolean tag values become strings — query `"true"`, not `true`.
- `targetQuery` is retroactive: tag a new resource and matching policies cover it automatically — a scope to watch when granting permissions.
- `memberQuery` ignores service accounts.
## Related
**access-control** (policy `targetQuery` / group `memberQuery` in context) · **cpln** (CLI command surface).
setup-agent7.8 KB
---
name: setup-agent
description: Deploys a Control Plane wormhole agent connecting workloads to private-network resources. Use when the user asks to reach a VPC, on-prem, data-center, or cross-cloud host, set up a tunnel, or run an agent.
---
# Agent Setup
A wormhole agent is a lightweight VM or container you run **inside the target network**. It opens a persistent **outbound** connection to Control Plane and tunnels workload traffic to any TCP/UDP endpoint on the private side — VPC, on-prem, data center, Azure VNet, cross-cloud, or a laptop. A workload reaches the endpoint by attaching an **identity** (gvc-scoped) that carries a `networkResources` entry pointing at the agent. No external egress firewall rule is needed.
> **Scope:** this is the deploy walkthrough — create the agent, deploy it on your platform, wire up the identity, verify. For the PrivateLink/PSC-vs-agent comparison, agent **sizing tables**, the full identity schema, and agent permissions, read **native-networking**. For the cloud-credential side (AWS/GCP/Azure access without an agent), read **setup-cloud-access**.
## Before you start
Confirm with the user: what private resource the workload must reach (host/IP + ports), where it lives (cloud/VPC/on-prem/cluster), the org, and whether an agent already exists (`list_resources` kind="agent"). **Reach for an agent only when PrivateLink/PSC does not fit** — for an AWS or GCP managed service, native networking is lower-latency and needs no agent (see native-networking). An agent is right for on-prem, cross-cloud, Azure, or local development. **Check which direction the user means:** an agent lets a Control Plane workload *reach into* their network; it does not run the workload on their hardware. For that — bare metal, on-prem VMs, a data-center server as a deployment target — the answer is a BYOK location (`mk8s-byok`).
## Step 1 — Create the agent
If the user asked you to set one up, create it directly; only list first when they want to reuse an existing one. Call `create_agent` (`org`, `name`, optional `description` / `tags`). The response contains the **bootstrap config JSON** — copy it out immediately.
CLI fallback (pipes the bootstrap straight to a file):
```bash
cpln agent create --name AGENT --org ORG > AGENT-bootstrap.json
```
> **Save the bootstrap config now.** It holds the registration token and is shown **only once, at creation**. Reads (`get_resource` kind="agent") return it with the token hidden. It is immutable — if lost, delete and recreate the agent. `update_agent` changes description / tags only.
## Step 2 — Deploy the agent
Pick the target; each artifact path is CLI/console (no MCP equivalent). Deploy in the **same VPC/region** as the target, with **outbound internet** and **no inbound ports** required.
| Target | How |
|---|---|
| **Kubernetes** | `cpln agent manifest --bootstrap-file AGENT-bootstrap.json -n NAMESPACE --replicas 2 > agent.yaml` then `kubectl apply -f agent.yaml`. Each agent stores a generated keypair as a K8s secret, so its service account needs secret create/modify in that namespace — use a dedicated namespace if that is a concern. |
| **Docker** (laptop / private host) | `cpln agent up --bootstrap-file AGENT-bootstrap.json` (one command, no manifest). `-b` runs it in the background; `--net` picks the Docker network. On Windows, disable the WSL 2 engine and run from a Windows prompt. |
| **AWS VM** | Subscribe to the **Control Plane Secure Communications Agent** in AWS Marketplace, launch via EC2 in the target VPC, enable a public IP or NAT for egress, and paste the bootstrap JSON into **User data**. Add the agent's security group to the target resource's inbound rules. |
| **Azure VM** | Azure Marketplace **Control Plane Secure Communications Agent** (gen-1); Public IP **None**, inbound **None**; paste the bootstrap JSON into **Custom data**. |
| **GCP VM** | `gcloud compute instances create … --metadata-from-file=user-data=AGENT-bootstrap.json` with the Control Plane agent image; open egress only (no SSH/RDP/ICMP needed). |
> **Never run two replicas of one deployment.** Each deployment has a unique key; duplicating it causes intermittent latency and dropped packets. For HA, run **separate** deployments (K8s `--replicas 2`; cloud VMs in a fixed-size instance group / ASG / VMSS). Agents run **active-active** — every instance serves traffic, and a missed-heartbeat instance is dropped while the group replaces it. The agent is **not CPU-intensive — do not autoscale on CPU**; size the group min 2, max = number of availability zones.
The agent also exposes a proxy on port **3128** (`cpln agent up --exposeProxy`) so systems inside the private network can call Control Plane workloads without external firewall changes — grant it on the workload's **Internal** firewall (see firewall-networking).
## Step 3 — Wire the identity to the agent
A workload routes through the agent only when an identity carrying a `networkResources` entry is attached to it. Create the identity first if it does not exist (`create_identity`, see access-control).
Add the agent-based resource with `add_identity_network_resource` (`org`, `gvc`, `identity`, one `resource`):
```json
{
"org": "ORG", "gvc": "GVC", "identity": "IDENTITY",
"resource": {
"name": "on-prem-db",
"agentLink": "//agent/AGENT",
"IPs": ["10.0.1.50"],
"ports": [5432]
}
}
```
Key constraints (Joi-enforced; mirrored by the tool): `name` unique across **both** `networkResources` and `nativeNetworkResources` and never equal to a FQDN; `IPs` (1-5 IPv4) **xor** `FQDN` (exactly one); `ports` 1-10, each 0-65535; optional `resolverIP` for private DNS; max 50 per array. `update_identity` replaces the whole array; `remove_identity_network_resource` deletes by name from either array (destructive — confirm first).
> For a local Docker agent, set the resource `IPs` to the host's **Docker network-adapter IP**, never `localhost` / `127.0.0.1`.
**Attach the identity to the workload:** `update_workload` setting `spec.identityLink` to `//identity/IDENTITY`. Without the attachment, nothing routes.
## Step 4 — Verify
- `get_agent_info` — `lastActive` within 60s means active; check `peerCount` and `serviceCount`. `get_agent_eventlog` shows connection events and errors. (CLI: `cpln agent info|eventlog AGENT --org ORG`.)
- `list_identity_network_resources` confirms the entry is on the identity.
- From the workload, dial the resource **`name`**. Use the `cpln` CLI after reading the `cpln` skill when an in-container connectivity probe is required. For a **TLS** target, connect on the **FQDN**, not the `name` — the certificate is issued for the FQDN.
## Common mistakes
- **Wrong deploy command** — `cpln agent up` is Docker hosts; `cpln agent manifest` is K8s; cloud VMs use the marketplace image + bootstrap as user-data (no `cpln` command).
- **Losing the bootstrap config** — output once at creation; if lost, delete and recreate.
- **Scaling one deployment past 1 replica** — drops packets; use separate deployments.
- **Missing `agentLink`, or `localhost` for a local agent's IP** — traffic cannot route.
- **Forgetting `spec.identityLink`** — the identity is wired but never reaches the workload.
- **Using `name` instead of `FQDN` for a TLS endpoint** — certificate validation fails.
## Related skills
| Need | Skill |
|---|---|
| PrivateLink/PSC vs agent, sizing, full identity schema, permissions | `native-networking` |
| Credential-free AWS / GCP / Azure access (no agent) | `setup-cloud-access` |
| The Internal firewall for the 3128 proxy, service-to-service rules | `firewall-networking` |
| Creating the identity, policies on the agent | `access-control` |
## Documentation
- [Agent Reference](https://docs.controlplane.com/reference/agent.md) · [Agent Setup Guide](https://docs.controlplane.com/guides/agent.md)
- [Identity Reference](https://docs.controlplane.com/reference/identity.md)
setup-cloud-access8 KB
---
name: setup-cloud-access
description: Credential-free cloud access (Universal Cloud Identity) for a Control Plane workload. Use when a workload needs AWS, GCP, Azure, or NATS NGS resources without embedded keys, or asks to register a cloud account.
---
# Cloud Access Setup (Universal Cloud Identity)
A workload reads cloud resources with **no embedded keys**: a GVC-scoped **identity** carries a per-provider cloud-access block that federates with the provider's IAM, and Control Plane vends short-lived credentials at runtime. Cloud SDKs (boto3, google-cloud, @azure/sdk) pick them up automatically — no SDK config.
## The chain
| Step | What must be true | Without it |
|---|---|---|
| **1. Cloud account** | a `cloud_account` (org-wide) maps to the provider, registered after the cloud-side IAM setup | identity can't federate |
| **2. Identity cloud block** | the identity carries an `aws`/`gcp`/`azure`/`ngs` block linking that cloud account | no credentials vended |
| **3. Workload link** | the identity is attached to the workload (`spec.identityLink`) | workload has no identity |
Order is strict: the cloud account must exist **before** the identity's cloud block references it.
## Key constraints
- **Identities are GVC-scoped** — one per workload, shareable within a GVC, never across GVCs. Same access in another GVC = recreate the identity there.
- **One cloud account per provider per identity** — one AWS + one GCP + one Azure + one NGS is fine; two AWS on one identity is not.
- **Cloud accounts are org-scoped** — always pass `org`.
- **Provider is immutable** — to switch providers, delete and recreate the cloud account.
## Step 1 — Cloud-side IAM setup
Each provider needs IAM configured **on the provider side first** so Control Plane can assume a role / impersonate a service account. Run the per-provider how-to to get the org-specific values (Control Plane's AWS account ID + external ID, the GCP service-account email, the Azure Function-App connector steps) — **never guess these**:
- `how_to_create_aws_cloud_account` — trust policy, account ID, external ID, the IAM permissions for the `cpln-connector` policy. Create an IAM role with that trust policy + connector policy + `ReadOnlyAccess`; note the **role ARN**.
- `how_to_create_gcp_cloud_account` — add the shown service account as an IAM principal with **Viewer, Project IAM Admin, Service Account Admin, Service Account Token Creator** (plus the service Admin role, e.g. `roles/storage.admin`, for each resource type identities will use); note the **project ID**.
- `how_to_create_azure_cloud_account` — create a Function App, deploy the connector, make it subscription **Owner**, capture the Function URL + `iam-broker` code into an `azure-connector` secret.
- `how_to_create_ngs_cloud_account` — create a `nats-account` secret holding your NATS account credentials.
CLI fallback: `cpln cloudaccount create-<provider> --how --org ORG`.
## Step 2 — Register the cloud account
`create_cloud_account` (`provider` = `aws`/`gcp`/`azure`/`ngs`), passing the value the provider needs:
| Provider | Required field |
|---|---|
| aws | `roleArn` (the role ARN from step 1) |
| gcp | `projectId` |
| azure | `secretLink` to an existing `azure-connector` secret (created by the user) |
| ngs | `secretLink` to an existing `nats-account` secret (created by the user) |
`status.usable` stays `false` until the cloud-side IAM exists. `update_cloud_account` edits the data block / tags (provider stays immutable). CLI fallback: `cpln cloudaccount create-aws|create-gcp|create-azure|create-ngs`.
## Step 3 — Identity with a cloud-access block
`create_identity` (or `update_identity` on an existing one) accepts the per-provider block directly — pass `aws`, `gcp`, `azure`, or `ngs`. On update each block **replaces wholesale**; `removeCloudIdentities: ["aws"]` detaches one. Every block needs `cloudAccountLink: //cloudaccount/NAME`. The pattern per provider:
- **aws** — Control Plane creates a new IAM role with `policyRefs` (managed = `aws::AmazonS3ReadOnlyAccess`, custom = bare name; chars `a-zA-Z0-9/+=,.@_-` only — **never full ARNs**), **xor** `roleName` to reuse a role. Optional `trustPolicy` (only alongside `policyRefs`).
- **gcp** — creates a service account with `bindings` (`resource` + `roles` like `roles/storage.objectViewer`; omit `resource` = project), **xor** `serviceAccount` to reuse one. Optional `scopes`.
- **azure** — creates a managed identity with `roleAssignments` (`scope` + `roles`; omit `scope` = subscription).
- **ngs** — scoped NATS creds: `pub`/`sub` `allow`/`deny` subjects (`*` single, `>` multi-level), `resp.max`/`resp.ttl`, and `subs`/`data`/`payload` limits (`-1` = no limit).
```yaml
spec:
aws:
cloudAccountLink: //cloudaccount/my-aws
policyRefs: ["aws::AmazonS3ReadOnlyAccess", "MyCustomPolicy"]
```
CLI fallback (MCP unavailable / CI-CD): `cpln identity get NAME --gvc GVC -o yaml-slim > id.yaml`, add the block under `spec`, `cpln apply -f id.yaml`. Confirm the exact shape with `get_resource_schema` (kind `identity`) before authoring YAML by hand.
## Step 4 — Link to the workload and verify
`update_workload` sets `spec.identityLink = //identity/NAME` (CLI: `cpln workload update NAME --set spec.identityLink=//identity/NAME`).
Read the identity back with `get_resource` (kind `identity`): `status.<provider>.usable` must be `true`; if `false`, read `status.<provider>.lastError`. Cloud CLIs are usually absent from production containers, so a missing `aws`/`gcloud`/`az` does **not** mean access is broken — the SDK path still works.
## Private-network resources
Reaching a private VPC / on-prem endpoint is a different mechanism on the **same identity**: a `networkResources` (agent/wormhole) or `nativeNetworkResources` (AWS PrivateLink / GCP PSC) array, not a cloud-access block. For the agent deployment walkthrough use **setup-agent**; for the comparison, producer-side setup, and the resource schema use **native-networking**.
## Common mistakes
- **Cloud block before the cloud account** — register the account first; the link won't resolve otherwise.
- **Skipping the how-to** — the org-specific account/external IDs and SA email are required and can't be guessed.
- **Full ARN in AWS `policyRefs`** — use the policy name with an optional `aws::` prefix, no colons.
- **Both `policyRefs` + `roleName` (AWS) or `bindings` + `serviceAccount` (GCP)** — exactly one.
- **Not checking `status.<provider>.usable`** — verify `true` before linking to the workload.
- **Confusing cloud access with secret access** — cloud access is the identity's `aws`/`gcp`/`azure`/`ngs` block; secret access is a `reveal` policy on a `cpln://secret/` reference (see **setup-secret**).
- **Sharing an identity across GVCs** — recreate it per GVC.
## Quick reference — MCP tools
| Tool | Purpose |
|---|---|
| `how_to_create_<provider>_cloud_account` | Org-specific cloud-side IAM steps (run first) |
| `create_cloud_account` / `update_cloud_account` | Register / edit a cloud account (provider immutable) |
| `get_resource` (kind `secret`) | Verify the NGS / Azure connector secret exists before referencing it |
| `create_identity` / `update_identity` | Create / edit the identity, including its cloud-access block |
| `update_workload` | Set `spec.identityLink` |
| `get_resource` / `list_resources` / `delete_resource` (kind `cloud_account` / `identity`) | Read / delete on any profile |
## Related skills
| Need | Skill |
|---|---|
| Private-VPC / on-prem connectivity, PrivateLink/PSC schema | native-networking |
| Deploy the wormhole agent for a private network | setup-agent |
| Identity, policy, and `reveal` for `cpln://secret/` refs | setup-secret |
| Policy shape, permissions, principals | access-control |
## Documentation
- [Accessing Cloud Resources](https://docs.controlplane.com/core/accessing-cloud-resources.md)
- [Create a Cloud Account](https://docs.controlplane.com/guides/create-cloud-account.md)
- [Cloud Account Reference](https://docs.controlplane.com/reference/cloudaccount.md) · [Identity Reference](https://docs.controlplane.com/reference/identity.md)
setup-secret8.54 KB
---
name: setup-secret
description: Secret access wiring and manifest authoring. Use when a workload needs to read a secret, the user asks to create a secret or generate secret YAML, configure a pull secret, or fix a deployment paused on a secret reference.
---
# Secret Access Setup
> **Secrets are read-only through this app.** `list_resources` / `get_resource` (kind="secret") show existence and metadata, never values. No tool creates, edits, deletes, or reveals a secret: secret data and lifecycle are managed by the user (Console, CLI, Terraform, Pulumi, or the API); you draft manifests with placeholders, the user fills the values and applies. `grant_workload_secret_access` grants a workload access — it never returns values.
Secret access is the #1 thing users get wrong: a workload reads a secret only when **three** things are all in place. Miss any one and the value is silently absent at runtime — or the deployment pauses on an unresolved reference.
## The mandatory chain
| Step | What must be true | Without it |
|---|---|---|
| **1. Identity** | an identity exists and is linked to the workload (`spec.identityLink`) | workload has no API credential — reads nothing |
| **2. Policy** | a policy grants that identity `reveal` on the secret | reference resolves to empty |
| **3. Reference** | the secret is injected as `cpln://secret/NAME` (env or volume) | nothing to read |
`reveal`, **not** `view` — `view` exposes only metadata. This is the single most common mistake.
## Pull secrets are different — no identity/policy
To pull images from a private registry, don't build the chain. Add the registry secret to the **GVC's** `pullSecretLinks` and every workload in that GVC can pull. Pull secrets are registry credentials — `docker`, `ecr`, or `gcp` types.
```yaml
kind: gvc
spec:
pullSecretLinks:
- //secret/my-registry
```
## Authoring a secret manifest — the user applies it
Drafting the manifest is an expected part of the job — users ask for a scaffold, fill in the real values themselves, and apply it. Generate the YAML with UPPERCASE placeholders, then always hand back the next steps:
1. **Fill in the placeholders locally** — the value never enters the chat.
2. **Apply it**: `cpln apply -f secret.yaml --org ORG`, the Console's **cpln apply** button (paste the YAML), or the per-type CLI command that reads the value from a file (`cpln secret create-docker --name NAME --file config.json`).
3. **Treat the filled file as a live credential** — keep it out of git and delete it after applying.
4. **Say when it's done** — verify with `get_resource` (kind="secret") and continue with the access chain below.
Never ask for the real value in chat, and never apply the manifest yourself.
`data` has a fixed shape per `type`, validated by the backend on create. The trap: **for `docker`, `gcp`, and `azure-sdk`, `data` is a single JSON string** (a `>-` block scalar in YAML), never a YAML mapping — an object is rejected.
```yaml
kind: secret
name: my-registry
type: docker
data: >-
{"auths":{"REGISTRY_HOST":{"username":"USERNAME","password":"PASSWORD"}}}
```
| `type` | `data` | Backend validation |
|---|---|---|
| `opaque` | object `{payload, encoding?}` | `payload` valid base64 when `encoding: base64` (default `plain`) |
| `dictionary` | object of string values | keys match `[-._a-zA-Z0-9]+` |
| `userpass` | object `{username, password, encoding?}` | — |
| `tls` | object `{cert, key?, chain?}` | `cert` and `key` must be valid PEM |
| `keypair` | object `{secretKey, publicKey?, passphrase?}` | `secretKey` a valid PEM private key |
| `aws` | object `{accessKey, secretKey, roleArn?, externalId?}` | `accessKey` starts `AKIA`, `roleArn` starts `arn:` |
| `ecr` | aws fields + `repos` (1–20) | each `ACCOUNT_ID.dkr.ecr.REGION.amazonaws.com[/REPO]` |
| `azure-connector` | object `{url, code}` | `url` must be https |
| `nats-account` | object `{accountId, privateKey}` | `accountId` a public nkey (`A…`), `privateKey` a seed (`SA…`) |
| `docker` | **JSON string** | must parse with an `auths` object keyed by registry host, at least one entry |
| `gcp` | **JSON string** | full service-account key: `type`, `project_id`, `private_key_id`, `private_key`, `client_email`, `client_id`, `auth_uri`, `token_uri`, `auth_provider_x509_cert_url`, `client_x509_cert_url` |
| `azure-sdk` | **JSON string** | `subscriptionId` / `tenantId` / `clientId` (UUIDs) plus `clientSecret` |
`get_resource_schema` (kind="secret") returns the apply schema and REST endpoints.
## Workflow
### 1 — Identify the secret
The secret must already exist — the user creates and rotates it through any Control Plane surface: Console, CLI (value-in-a-file, never an inline flag), Terraform, Pulumi, or the API. Confirm it exists with `list_resources` or `get_resource` (kind="secret") before wiring anything; never ask for the value in chat and never invent a placeholder. If it does not exist yet, author the manifest (section above) and wait until the user has applied it.
### 2 — Grant the workload access
**Preferred — one call.** `grant_workload_secret_access` (`gvc`, `workloadName`, `secretName`) creates the identity if missing (default `{gvc}-{workloadName}`), links it to the workload, and creates/updates a `reveal` policy (default `{gvc}-{workloadName}-secrets-policy`). It never returns secret values, and it does **not** inject the reference — step 3 still applies.
**Manual alternative** (granular control): `create_identity` → `update_workload` to set `spec.identityLink` → `create_policy` (targetKind `secret`, a `reveal` binding naming the identity). Policy shape lives in **access-control**.
**Ordering matters.** The workload must already exist. For a new workload that references a secret: `create_workload` first (its deployment pauses on the unresolved reference), then grant — the deployment resumes.
Identities are **GVC-scoped**: one per workload, shareable across workloads in the same GVC, never across GVCs.
### 3 — Inject the reference
`update_workload` (read current state with `get_resource` first) to add `cpln://secret/NAME` — the whole secret — or `cpln://secret/NAME.KEY` for one property:
| Type | Keys | Example |
|---|---|---|
| opaque | `payload` | `cpln://secret/api-key.payload` |
| userpass | `username`, `password` | `cpln://secret/creds.password` |
| tls | `key`, `cert`, `chain` | `cpln://secret/web-tls.cert` |
| dictionary | user-defined | `cpln://secret/cfg.DB_HOST` |
| aws / ecr | `accessKey`, `secretKey`, `roleArn` | `cpln://secret/aws.accessKey` |
Inject as an **env var** or a **volume mount** (`{ uri: "cpln://secret/NAME", path: "/secrets/x" }`). Mounts are read-only (except Azure Files), max **15** per container, and these knative-reserved paths are rejected: `/dev`, `/dev/log`, `/tmp`, `/var`, `/var/log`.
### 4 — Verify and redeploy
- `get_resource` (kind="workload") → `spec.identityLink` is set and the env/volume reference reads `cpln://secret/…`.
- `get_resource` (kind="policy") → the binding grants `reveal` to that identity.
- Updating a workload spec redeploys automatically; via CLI use `cpln apply --ready` to block until healthy. **A rotated secret value needs a redeploy** — running replicas keep the old value until then.
## Quick reference — MCP tools
| Tool | Purpose |
|---|---|
| `grant_workload_secret_access` | Composite — identity + `reveal` policy + link, in one call |
| `create_identity` / `create_policy` | Build the access chain manually (granular control) |
| `update_workload` | Set `identityLink`; inject the env / volume reference |
| `list_resources` / `get_resource` (kind="secret") | Confirm a secret exists / read its metadata |
## Common mistakes
- **Object `data` on a docker / gcp / azure-sdk secret** — those three types take one JSON string; a YAML mapping fails validation.
- **No identity** — a workload with no `identityLink` reads no secrets.
- **`view` instead of `reveal`** — metadata only, no value.
- **Bad reference** — must be `cpln://secret/NAME`, not the bare name.
- **Granting before the workload exists** — the workload comes first.
- **Sharing an identity across GVCs** — they are GVC-scoped.
- **Over-engineering pull secrets** — registries need only `pullSecretLinks`, no identity/policy.
- **Skipping the redeploy after rotation** — running replicas keep the old value.
## Related skills
| Need | Skill |
|---|---|
| Policy shape, permissions, principals | `access-control` |
| Workload identities, cloud / private-network access | `native-networking` |
| Workload spec, deploy, env vars | `workload` |
## Documentation
- [Secret Reference](https://docs.controlplane.com/reference/secret.md)
stateful-storage11.9 KB
---
name: stateful-storage
description: "Creates persistent storage for stateful workloads on Control Plane. Use when the user asks about volumes, volume sets, disks, mounting storage, snapshots, volume expansion, filesystems, shared storage, or backups."
---
# Stateful Storage & VolumeSets
A **VolumeSet** is GVC-scoped persistent storage for workloads. The `workload` skill covers the basics (stateful type, reserved mount paths, the 15-volume limit, create-then-verify); this skill is the full volume-set detail. The one trap that drives most rework: **`fileSystemType` and `performanceClass` are immutable** (a PATCH that changes either returns HTTP 400) — to change either you must create a new volumeset, and the old data does not carry over. Choose both at creation.
**Most databases don't need this skill:** `template-catalog` installs Postgres, Redis, MySQL, MongoDB, and more with the volumeset, snapshots, and credentials already wired — hand-build only for a custom app or an unsupported engine.
## Filesystem types and performance classes
| Filesystem | Access | Workloads | Volumes provisioned | Snapshots / shrink / delete-volume |
|---|---|---|---|---|
| `ext4` | read-write-once | one stateful/vm workload | one per replica, per location | yes |
| `xfs` | read-write-once | one stateful/vm workload | one per replica, per location | yes |
| `shared` | read-write-many | any workload type, many at once | one per location (shared by all replicas) | no — expand only |
| Performance class | Min | Max | Filesystems |
|---|---|---|---|
| `general-purpose-ssd` | 10 GB | 65536 GB | ext4, xfs |
| `high-throughput-ssd` | 200 GB | 65536 GB | ext4, xfs |
| `shared` | 10 GB | 65536 GB | shared (auto-set) |
When `fileSystemType: shared`, `performanceClass` is **auto-set to `shared`** — do not specify another. Data is **per-location** and never replicated across locations; for cross-location redundancy, replicate at the application layer (e.g. WAL streaming).
## Create a volumeset
Use `create_volumeset` (MCP create/mount tools default `fileSystemType` to **xfs** and `performanceClass` to `general-purpose-ssd`; the raw API/`cpln apply` default is **ext4**). YAML for IaC / CLI fallback:
```yaml
kind: volumeset
name: pg-data
gvc: GVC
spec:
fileSystemType: ext4
performanceClass: general-purpose-ssd
initialCapacity: 20 # GB; within the class min/max and <= autoscaling.maxCapacity
autoscaling:
maxCapacity: 100
minFreePercentage: 20 # 1-100
scalingFactor: 1.5 # >= 1.1
snapshots:
schedule: "0 2 * * *" # cron; no more than once per hour
retentionDuration: 7d # float + d/h/m; tool default 7d
```
Apply with `cpln apply -f volumeset.yaml --gvc GVC`. Update mutable fields (capacity, autoscaling, snapshot policy, tags) with `update_volumeset`.
### Autoscaling
**Reactive**: a background job checks volumes about once a minute; when free space falls below `minFreePercentage` it resizes the volume to hold current usage at that margin, scaled up: `new_capacity = ceil(usedGB / (1 - minFreePercentage/100) * scalingFactor)`, capped at `maxCapacity`. Both fields are required, or autoscaling does nothing.
**Predictive** runs the same formula on *projected* usage (from the recent growth rate) to expand ahead of demand; the larger of the reactive and predictive targets wins. Requires `minFreePercentage > 0` and `scalingFactor >= 1.1`:
```yaml
autoscaling:
maxCapacity: 200
minFreePercentage: 20
scalingFactor: 1.5
predictive:
enabled: true # default false
lookbackHours: 24 # 1-168
projectionHours: 6 # 1-72
minDataPoints: 10 # 2-100
minGrowthRateGBPerHour: 0.01
scalingFactor: 1.2 # >= 1.1; defaults to the parent scalingFactor
```
## Mount to a workload
Mount with `mount_volumeset_to_workload` — it attaches to the **first container** and creates the volumeset if missing (create-only defaults, ignored when the volumeset already exists: path `/mnt/{volumesetName}`, filesystem `xfs`, class `general-purpose-ssd`). Volume URI is `cpln://volumeset/VOLUMESET`.
- **ext4/xfs require a `stateful` or `vm` workload** (mounting on serverless/standard returns HTTP 400); `shared` mounts on any type. Workload type is immutable — see "Migrating to stateful" below.
- Up to **15 volumes** per container. **Reserved mount paths** (rejected): `/dev`, `/dev/log`, `/tmp`, `/var`, `/var/log`.
- `recoveryPolicy`: `retain` (default — reuse an existing volume's data on a new replica) or `recycle` (start fresh).
- `path` is required for non-vm workloads and rejected for `vm` (VM disks use `name`/`bus`/`bootOrder` instead).
- Stateful workloads give each replica a stable index and its own volume; `spec.loadBalancer.replicaDirect` (stateful-only) exposes per-replica endpoints — see the `workload` skill.
```yaml
kind: workload
name: pg
gvc: GVC
spec:
type: stateful
containers:
- name: postgres
image: //image/postgres:16
ports:
- number: 5432
protocol: tcp # http | http2 | grpc | tcp — a DB is tcp, not http
volumes:
- uri: cpln://volumeset/pg-data
path: /var/lib/postgresql/data
```
## Snapshots
Snapshots are **ext4/xfs only — never `shared`**. Automatic policy lives in `spec.snapshots`: `createFinalSnapshot` (default `true` — a snapshot is taken before any volume in the set is deleted), `retentionDuration`, and `schedule` (cron whose minute field must be a single concrete value, so no more than once per hour). Manual: `create_volumeset_snapshot`, `list_volumeset_snapshots`, `restore_volumeset_snapshot`, `delete_volumeset_snapshot`. A restore creates a **new volume** and discards everything written since the snapshot.
## Resize and delete volumes
- **Expand** — live, no downtime, all filesystems. Throttled to **4 expansions per volume per rolling 24 hours**; the 5th returns **HTTP 429** and a brief wait does not help (the oldest expansion must age out of the window). `expand_volumeset`.
- **Shrink** (ext4/xfs only) — data is migrated to the new smaller volume via an online presync + final delta sync, and the replica restarts during the swap. The platform **rejects the shrink with HTTP 400 when known used bytes (+5% metadata headroom) would not fit**; data is only lost if used bytes genuinely exceed the new capacity. Floor is the class minimum (10 / 200 GB). `shrink_volumeset`.
- **Delete a volume** (ext4/xfs only) — permanent loss of that volume's data. `delete_volumeset_volume`.
Shrink, volume-delete, snapshot-delete, and restore are destructive: **snapshot first** as the recovery net, then confirm the blast radius (the destructive-ops guardrail returns an impact preview before executing).
## Shared filesystem
A `shared` volumeset is mounted read-write by many workloads at once but supports only expand — no snapshots, shrink, or volume-delete. Each mount point is provisioned its own CPU/memory; tune with `mountOptions.resources` (defaults `minCpu 500m`, `maxCpu 2000m`, `minMemory 1Gi`, `maxMemory 2Gi`; max/min at most 4000m and 4096Mi apart, ratio at most 4:1):
```yaml
spec:
fileSystemType: shared # performanceClass auto-set to "shared"
initialCapacity: 50
mountOptions:
resources: { minCpu: 500m, maxCpu: 2000m, minMemory: 1Gi, maxMemory: 2Gi }
```
## Custom encryption (AWS only)
Volumes are encrypted by default. To use your own AWS KMS keys on ext4/xfs volumes (not `shared`, not BYOK):
```yaml
spec:
customEncryption:
regions:
aws-us-east-1: # format: {cloud-provider}-{region}
keyId: "arn:aws:kms:us-east-1:123456789:key/KEY_ID"
```
The `keyId` is injected as the EBS storage-class `kmsKeyId`. The KMS key policy **must grant Control Plane's AWS account `arn:aws:iam::957753459089:root`** the volume-encryption permissions (`Decrypt`, `Encrypt`, `GenerateDataKey`, `CreateGrant`, etc.); the key is immutable once a volume exists.
## BYOK storage classes
On self-hosted clusters, volumes use the storage class `{performanceClass}-{fileSystemType}` (e.g. `general-purpose-ssd-ext4`) and the cluster needs a CSI-compatible driver. `spec.storageClassSuffix` selects an alternative `{performanceClass}-{fileSystemType}-{suffix}`, falling back to the unsuffixed class if it is not found.
## Migrating a workload to stateful
Workload type is immutable, so adding an ext4/xfs volume to a serverless/standard workload means **delete + recreate as `stateful`** — destructive. Before deleting, confirm with the user: the public URL that 5xx's during the cutover, any internal callers that fail until recreate, runtime/in-memory state lost at delete, and that the recreate typically takes 2-5 min. Sequence: capture the spec (`cpln workload get WORKLOAD --gvc GVC -o yaml-slim > bak.yaml`) as a rollback artifact; apply the volumeset; delete the old workload; apply the new manifest with `spec.type: stateful` + the volume mount, **keeping the same name** to preserve URL/DNS/policy/identity links. For the deploy-wait pattern, see the `workload` skill.
## Verify
- `get_resource` (kind `volumeset`): `status.locations[].volumes[]` show per-volume `currentSize`, `currentBytesUsed`, `lifecycle` (expect `bound`), and snapshot counts; `status.usedByWorkload` names the bound workload.
- After mounting, poll `list_deployments` until ready and confirm the container's volume is mounted at the expected path.
## Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| 400 mounting a volumeset | ext4/xfs on a serverless/standard workload | Use a `stateful` or `vm` workload (recreate to change type) |
| 400 "performanceClass / fileSystemType is immutable" | Tried to change either on update | Create a new volumeset; migrate data via snapshot/restore |
| 400 on mount with a path | Path is reserved (`/dev`, `/tmp`, `/var`, ...) | Mount elsewhere (e.g. `/data`, `/mnt/...`) |
| HTTP 429 on expand | 4 expansions on that volume in the last 24 h | Wait for the oldest to age out; plan larger steps |
| 400 on shrink | New size cannot hold used bytes (+5%) | Shrink less, or free space / snapshot then rebuild |
| Snapshot fields rejected | Volumeset is `shared` | Snapshots need ext4/xfs |
| Deployment stuck after mount | Volume provisioning (2-5 min on first deploy) | Poll `list_deployments`; check `get_workload_logs` if it stays unready |
## MCP tools quick reference
| Tool | Purpose | Tier |
|---|---|---|
| `create_volumeset` | Create a volumeset | core |
| `update_volumeset` | Update mutable fields (capacity, autoscaling, snapshots, tags) | core |
| `mount_volumeset_to_workload` | Mount to a workload (creates the volumeset if missing) | core |
| `expand_volumeset` | Grow a volume (4 / 24 h limit) | core |
| `shrink_volumeset` | Shrink a volume (ext4/xfs) | full |
| `delete_volumeset_volume` | Delete one volume (ext4/xfs) | full |
| `create_volumeset_snapshot` | Point-in-time snapshot | full |
| `list_volumeset_snapshots` | List snapshots | full |
| `restore_volumeset_snapshot` | Restore a snapshot to a new volume | full |
| `delete_volumeset_snapshot` | Delete a snapshot | full |
| `get_resource` / `list_resources` / `delete_resource` (kind `volumeset`) | Read / list / delete a volumeset | core |
CLI fallback (CI/CD via a service-account `CPLN_TOKEN`): `cpln volumeset create|get|update|delete|expand|shrink`, `cpln volumeset snapshot create|get|restore|delete`, `cpln volumeset volume get|delete`; `expand`/`shrink` need `--new-size` (`--location`/`--volume-index` optional), and `cpln apply -f` for YAML.
## Related skills
| Skill | For |
|---|---|
| `workload` | Workload types, the deploy-and-verify flow, load-balancer/`replicaDirect` config |
| `template-catalog` | Postgres, Redis, and other databases that provision volumesets for you |
| `firewall-networking` | Outbound rules for cloud-bucket volumes (`s3://`, `gs://`, `azureblob://`) |
## Documentation
- [Volume Set Reference](https://docs.controlplane.com/reference/volumeset.md)
- [Workload Volumes](https://docs.controlplane.com/reference/workload/volumes.md)
- [CLI volumeset Commands](https://docs.controlplane.com/cli-reference/commands/volumeset.md)
tag9.14 KB
--- name: tag description: "Resource tags on Control Plane — labels that organize resources, trigger built-in cpln/ behaviors, and drive targeting. Use when the user asks about tags, labels, tagging, naming conventions, or resource protection." --- # Tags — Labeling, Built-in Behaviors & Targeting Tags are **key-value labels** on almost every Control Plane resource (workload, GVC, identity, secret, policy, group, domain, image, volumeset, agent, ipset, org, mk8s, user, service account). They live in a top-level `tags:` map. A tag is three things at once: **metadata** for humans, a **selector** that policies/groups/GVCs/queries target, and — under the reserved `cpln/` namespace — a **switch** that turns on platform behavior the spec doesn't yet expose as a field. ## Why tags pay off | Capability | What a tag unlocks | Mechanism | |:---|:---|:---| | Dynamic RBAC | A policy `targetQuery` on `environment=production` grants on every match — **including resources created later** | **access-control** | | Dynamic group membership | A group `memberQuery` auto-enrolls users by tag (e.g. SSO provider) | **access-control** | | Dynamic placement | A GVC `locationQuery` picks locations by tag (e.g. `cpln/country`) instead of a fixed list | **query-spec** | | Fleet inventory & bulk ops | `cpln KIND query --tag tier=frontend` finds every matching resource to act on | **cpln** | | Built-in behaviors | Reserved `cpln/*` tags switch on features (protection, sticky sessions, mTLS, …) | see below | | Console organization | List columns, custom logos, saved groups, and the Query filter all read tags | see below | The payoff is **retroactive and self-maintaining**: tag a new workload `environment=production` and every prod policy, group, and dashboard that queries that tag covers it automatically — no rule edits. ## Setting tags | Where | How | |:---|:---| | CLI, dedicated | `cpln KIND tag NAME --tag key=value` (repeatable); `--remove-tag key` drops one | | CLI, on create | `cpln KIND create ... --tag key=value` | | CLI, generic update | `cpln KIND update NAME --set tags.key=value` (kinds with no `tag` subcommand, e.g. `user`) | | MCP | The `create_*` / `update_*` tool for the kind accepts a `tags` object | | Manifest | A top-level `tags:` map, then `cpln apply -f FILE` | ```bash cpln workload tag my-api --tag environment=production --tag team=payments cpln workload tag my-api --remove-tag team ``` **`=` guesses the type, `:` forces a string.** `--tag replicas=3` stores the number `3`; `--tag replicas:3` stores the string `"3"`; an empty value (`--tag key=`) stores `null`. This matters because queries are type-sensitive (see Gotchas). ## A taxonomy that earns its keep Tags become leverage only when keys and values are **uniform** — a query for `environment=production` silently misses anything tagged `env=prod` or `Environment=Production`. Agree on a small, lowercase vocabulary up front: | Key | Example values | Drives | |:---|:---|:---| | `environment` | `production`, `staging`, `dev` | RBAC scope, dashboards, promotion | | `team` / `owner` | `payments`, `platform` | ownership, on-call routing, group queries | | `tier` | `frontend`, `backend`, `data` | fleet ops, firewall / policy scope | | `app` | `checkout`, `billing-api` | grouping multi-workload apps | | `managed-by` | `terraform`, `console` | drift detection, IaC ownership | Tag for the keys you will actually query; a tag nobody selects on is just decoration. Stay out of the `cpln/`, `syncer.cpln.io/`, and `firebase/` prefixes — those are platform-defined (below). ## Built-in tags that change behavior The `cpln/` namespace is reserved: don't invent your own keys under it, but **do** set the documented tags below to switch on behavior. They are the escape hatch for options not yet first-class fields. **Any resource — deletion guard.** `cpln/protected=true` makes the platform refuse to delete the resource (any kind); remove the tag to delete. The MCP `delete_resource` and `cpln KIND delete` both fail until it's cleared. In the Console it's the lock switch next to **Actions**. ```bash cpln workload tag WORKLOAD --tag cpln/protected=true # block delete cpln workload tag WORKLOAD --remove-tag cpln/protected # allow delete ``` **Workload behavior:** | Tag | Value | Effect | |:---|:---|:---| | `cpln/timeoutSecondsOverride` | seconds (≤3600) | Raise the request timeout past the 600s ceiling | | `cpln/largeDisk` | `true` | Allocate a large ephemeral disk | | `cpln/tracingDisabled` | `true` | Turn off distributed tracing for this workload | | `cpln/publishNotReadyAddresses` | `true` | Route internal traffic to replicas before they pass readiness | | `cpln/discoverCrossGvcReplicas` | `true` | Discover replicas in other GVCs over mTLS | | `cpln/bypassProxyOutbound` | `true` | Skip the service-mesh proxy on outbound traffic | | `cpln/disableServiceMeshInboundPort` / `...OutboundPort` | port | Exclude one port from the service mesh | | `cpln/externalAuth*` | family | Route every request through an external authorization service (`...Address` required; see **workload-security**) | | `cpln/rateLimit*` | family | Enforce limits via an external rate-limit service (`...Address` required; see **cdn-rate-limiting**) | BYOK / Direct-LB workloads add `cpln/disableServiceMesh`, `cpln/disableServiceMeshOutboundCIDR`, and `cpln/k8sClusterRole`. **GVC — sticky sessions** (apply to every workload in the GVC): `cpln/sessionCookie` (cookie name) plus `cpln/sessionDuration` (a Go duration, e.g. `30m`). **Domain:** `cpln/clientCertificateValidation=enabled` requires a valid client cert, i.e. mTLS (**domain**); `cpln/skipDNSCheck=true` skips DNS validation; `cpln/wildcard=true` enables a wildcard certificate. ## Auto-populated tags you can target Some tags are set *by the platform* and are read-only — their value is that you can **query** them: | Tag | On | Use | |:---|:---|:---| | `cpln/city` / `cpln/country` / `cpln/continent` | location | Select locations by geography in a GVC `locationQuery` | | `firebase/sign_in_provider` | user | Auto-enroll users by SSO provider in a group `memberQuery` | | `syncer.cpln.io/source` / `syncer.cpln.io/lastError` | secret | Trace External-Secret-Syncer ownership and its last sync error | | `cpln/release` | helm-managed resources | Identify what `cpln helm` created | ## In the Console | Tag | UI behavior | |:---|:---| | `cpln/protected` | Toggled by the lock switch; blocks the Delete action | | Any tag whose value is a URL (`https://`, `http://`, `ws://`, `wss://`) | Rendered as a clickable link in the resource's Tag Links | | `cpln/console.tagColumns` (org) | Surfaces chosen tags as list columns, e.g. `workload=env,team;gvc=region` | | `cpln/custom-logo` / `cpln/custom-logo-dark` (org) | Custom org logo in the sidebar | | `resourceGroup::KEY[::VALUE]` (org) | Saved, pinnable resource groups in the sidebar nav | The **Query** button on any list maps directly to the query spec (match All / Any / None) — see **query-spec**. ## Worked examples **Production RBAC by tag (retroactive).** Tag the workloads, then grant once: ```bash cpln workload tag checkout payments-api --tag environment=production cpln policy create --name prod-operators --target-kind workload \ --query-tag environment=production # then add a binding — see access-control ``` New workloads tagged `environment=production` fall under the policy automatically. **Fleet query.** Find every frontend workload to audit or roll: `cpln workload query --tag tier=frontend`. ## Gotchas | Trap | Detail | |:---|:---| | Booleans match as strings | A tag stored `true` is the string `"true"` — query `value: "true"`, not `true`, or it matches nothing | | Set-time type matters | `--tag n=5` stores a number, `--tag n:5` a string; query the same type you stored | | No inheritance | A GVC's tags do **not** flow to its workloads — tag each resource you want matched | | Mutable on immutable resources | Name and type are fixed, but tags are always editable — even on the org | | Protected blocks delete | `cpln/protected=true` makes deletes fail until you remove the tag | | Removal | `--remove-tag key` (or set the value `null`); an empty value is kept as a marker, not deleted | | Case-sensitive | Keys and values are exact-match; `Prod` is not `prod` | ## Verify - `cpln KIND get NAME -o yaml` — confirm the `tags:` block. - `cpln KIND query --tag key=value` — confirm the resource is selected by the tag a policy or group will use. ## Related skills - **query-spec** — the query language: operators, match modes, and the three fields tags feed (`targetQuery`, `memberQuery`, `locationQuery`). - **access-control** — policy `targetQuery` and group `memberQuery` in context. - **workload-security** / **cdn-rate-limiting** — the `cpln/externalAuth*` and `cpln/rateLimit*` tag families. - **domain** — `cpln/clientCertificateValidation`, `cpln/skipDNSCheck`, `cpln/wildcard`. - **cpln** — the full CLI resource-command map. ## Documentation - [Tags](https://docs.controlplane.com/core/misc.md) · [Query](https://docs.controlplane.com/core/query.md) · [Resource Protection](https://docs.controlplane.com/guides/resource-protection.md) · [Workload special tags](https://docs.controlplane.com/reference/workload/general.md)
template-catalog10 KB
--- name: template-catalog description: "Recommends and installs templates from the Control Plane Template Catalog. Use when the user wants postgres, redis, kafka, mongodb, mysql, or any database, cache, queue, or gateway, or asks what templates exist." --- # Template Catalog The Template Catalog ships production-tested charts (Helm under the hood) for databases, caches, queues, brokers, search, gateways, and more — persistent storage wired up, credentials generated as Control Plane secrets, a sane firewall posture, and HA variants where they matter. For any common component the catalog template is the **default recommendation, not the fallback**: hand-rolled workload + volumeset + secret + firewall stacks routinely ship without backups, with a public database, or single-replica. Lead with the template, and move to a custom workload only when the user has a hard reason — an unusual extension, a legacy image they must reuse, or a feature the template doesn't expose. Template-first is also enforced by the operating guide's skill router. ## Find the right template `browse_templates` returns the **live catalog** — name, category, latest version, a "creates its own GVC" flag, and description. It is the source of truth for what exists; the table below is only the common asks. Filter with a substring (e.g. `postgres`), then call `get_template <name>` for the version list, prerequisites, and an example `values.yaml` to copy. | Need | Templates | |---|---| | PostgreSQL | `postgres` (single + backup), `postgres-highly-available` (Patroni failover), `pgedge` (active-active multi-master), `postgis` (geospatial) | | MySQL-compatible | `mysql`, `mariadb`, `tidb` (distributed) | | Distributed SQL | `cockroach`, `tidb` | | Document / NoSQL | `mongodb` (single), `mongodb-cluster` (replica set), `cassandra` | | Analytics / columnar | `clickhouse` | | Cache / KV | `redis` (replica + Sentinel), `redis-cluster` (sharded), `redis-multi-location` (cross-GVC), `etcd` | | Streaming / queues | `kafka`, `redpanda`, `rabbitmq`, `nats`, `cpln-task-runner` | | Search / vector | `manticore`, `opensearch`, `elasticsearch`, `weaviate` | | Gateway / WAF / VPN | `nginx`, `tyk`, `coraza`, `tailscale` | | Storage / AI / LLM | `minio` (S3), `ollama`, `langfuse` | | Auth / dev / ops | `fusionauth`, `dbeaver`, `airflow`, `ess`, `secret-env-var-syncer`, `otel-collector` | ## Choosing an HA / scaling variant This is the choice the catalog can't make for you: - **Postgres:** `postgres` is one instance with optional scheduled S3/GCS backups; `postgres-highly-available` adds Patroni leader election and an embedded etcd quorum (odd member count — 3/5/7) plus its own scheduled backups (logical or WAL-G mode); `pgedge` is active-active multi-master across regions. Pick HA when failover matters, pgEdge when you need multi-region writes. - **MongoDB:** `mongodb` is single; `mongodb-cluster` is a replica set and creates its own GVC. - **Redis:** `redis` (master-replica + Sentinel) for one location; `redis-cluster` (sharded, needs 6+ nodes) for horizontal scale; `redis-multi-location` (Valkey + Sentinel) for cross-location failover. - **Distributed SQL:** `cockroach` and `tidb` are natively distributed — HA is built in through their consensus protocols, and they create their own multi-location GVCs. - **Streaming:** `kafka` for the full Kafka ecosystem; `redpanda` is Kafka-API-compatible with a simpler single-binary footprint. ## Install (MCP) 1. `get_template <name>` — copy the example `values.yaml`; edit credentials, replica count, resources, storage size, and access scope. 2. `preview_template` (full profile) — dry-run render the resources the install would create, without applying anything. 3. `install_template` — pass `org`, a unique `name` (the release id, immutable), `template`, the `values` YAML (required, max 128 KiB), an optional `version` (latest if omitted), and `gvc`. **Omit `gvc` for templates that create their own** (the `createsGvc` flag in `browse_templates` / `get_template`; e.g. `cockroach`, `tidb`, `nats`, `clickhouse`, `airflow`, `mongodb-cluster`, `redis-multi-location`, `pgedge`). 4. Installs are asynchronous — confirm with `get_installed_template <name>`. ## Configure and upgrade Reconfigure with `upgrade_template`: pass `name` plus the new `version` and/or `values`. **`values` REPLACES the release's values entirely — there is no reuse-merge** — so start from the current values, never a partial. `template` and `gvc` are immutable and read from the installed release, so you don't pass them. Roll back with `rollback_template` (full profile) or `cpln helm rollback`. Access scope lives in the template's `values` — but the **key name varies per template** (e.g. `internal_access.type`, `internalAccess.type`, `internalAllowType`, or `firewall.internal_inboundAllowType`). Values are `same-gvc` (default), `same-org`, `workload-list` (with an explicit `workloads:` list), and `none` on a few. Copy the example from `get_template` rather than writing keys from memory. ## CLI fallback (CI/CD) When MCP is unavailable, or in pipelines with a service-account `CPLN_TOKEN`, use `cpln helm` against the OCI registry `oci://ghcr.io/controlplane-com/templates/<TEMPLATE>` (the slug is the template name): ```bash cpln helm install my-pg oci://ghcr.io/controlplane-com/templates/postgres -f values.yaml # omit --version for latest cpln helm template my-pg oci://ghcr.io/controlplane-com/templates/postgres -f values.yaml # preview rendered resources cpln helm list # releases in the org cpln helm get values <RELEASE> --all # currently applied values cpln helm upgrade <RELEASE> oci://... -f values.yaml cpln helm history <RELEASE> # revision numbers, for rollback cpln helm rollback <RELEASE> [<REVISION>] # previous revision if omitted cpln helm uninstall <RELEASE> ``` Reference `values.yaml` for any template lives in the [templates repo](https://github.com/controlplane-com/templates) at `<template>/versions/<version>/values.yaml`. ## Verify After install, `get_installed_template <name>` shows the release status, revision, and every resource it created (it decodes the release secret, so the token needs secret **reveal** permission). Then confirm the workloads are healthy with `list_deployments`, and check the generated secrets with `list_resources` (kind `secret`). Add firewall rules or a domain for any workload that needs external access. **Connection details:** an installed service is reachable inside its GVC at `<release>-<component>.<gvc>.cpln.local:<port>` (e.g. `my-pg-postgres.<gvc>.cpln.local:5432`); credentials live in the generated dictionary secret — the user reads them in the Console; reference them in workloads as `cpln://secret/NAME.KEY`. The exact workload name, port, and secret keys are in `get_installed_template` (created resources) and the `get_template` example values. ## Traps - `upgrade_template` **replaces** values — there is no partial merge, so start from the current values, not a fragment. - Templates with `createsGvc` make their own GVC — **omit `gvc`** on install; others require an existing `gvc`. - Uninstall removes the resources the release created, **including volume data** — confirm the blast radius first. - `values` key names (credentials, resources, access scope) differ across templates — copy from `get_template`, don't hand-write from memory. - Backups (e.g. `postgres`, `mongodb`) need a **Cloud Account + storage IAM policy** to exist first, referenced in the `values` backup block — see the `get_template` prerequisites. ## Troubleshooting | Symptom | Cause | Fix | |---|---|---| | Install fails: GVC required | template installs into an existing GVC | pass `gvc` (or check `createsGvc` — self-GVC templates omit it) | | Upgrade lost settings | `values` replaces, not merges | re-supply full values from `get_template` / `cpln helm get values --all` | | `get_installed_template` permission denied | token lacks secret reveal | grant `reveal` on the release secret via a policy | | Install failed / release stuck | partial apply, bad values, or unready workloads | inspect `get_installed_template` and `cpln helm history`; fix values and `upgrade_template`, or `uninstall` and reinstall | | Workloads pending after install | image pull / firewall / resources | see the `workload` skill's troubleshooting | ## Quick reference | Tool | Purpose | Tier | |---|---|---| | `browse_templates` | Live catalog (filter by substring) | core | | `get_template` | Versions, prerequisites, example values | core | | `preview_template` | Dry-run render, no apply | full | | `install_template` | Install a release | core | | `upgrade_template` | Change version/values (replaces values) | core | | `rollback_template` | Roll back to a prior revision | full | | `uninstall_template` | Remove a release and its resources | core | | `list_installed_templates` | Inventory of releases in the org | core | | `get_installed_template` | One release's status + created resources | core | CLI fallback (CI/CD via a service-account `CPLN_TOKEN`): `cpln helm install|template|list|get|upgrade|rollback|uninstall|history` against `oci://ghcr.io/controlplane-com/templates/<TEMPLATE>`. ## Related skills | Skill | For | |---|---| | `workload` | Custom workloads when no template fits; deploy-and-verify; pending-replica troubleshooting | | `stateful-storage` | Volumesets the database templates provision; snapshots and expansion | | `firewall-networking` | Exposing an installed service; outbound rules | | `access-control` | Identities and policies for backup cloud accounts and secret reveal | | `iac-terraform-pulumi` | Installing templates through Terraform or Pulumi | ## Documentation - [Template Catalog Overview](https://docs.controlplane.com/template-catalog/overview.md) - [Install via CLI](https://docs.controlplane.com/template-catalog/install-manage/cli.md) · [Terraform](https://docs.controlplane.com/template-catalog/install-manage/terraform.md) · [Pulumi](https://docs.controlplane.com/template-catalog/install-manage/pulumi.md) · [UI](https://docs.controlplane.com/template-catalog/install-manage/ui.md)
workload24.6 KB
---
name: workload
description: "Primary skill for creating, updating, running, and debugging workloads on Control Plane; routes to a deeper skill per subject. Use when the user asks to deploy or run a container, app, API, service, worker, or job, or to change, scale, expose, secure, or diagnose one."
---
# Workloads — Primary Skill & Router
A **workload** is Control Plane's unit of deployment: one or more containers plus how they scale, get exposed, store data, and stay healthy. This skill carries the must-know primary rules for safely creating, updating, and running a workload.
**Need more detail on one subject?** This skill covers the common case; for depth on a single topic, load the matching skill from the **Deep-dive router** at the end — you may load one or several, as the task spans. The skill files are already loaded — open the relevant skill(s) directly.
## Workload type — the first decision (standard is the default)
`create_workload` defaults the type to **`standard`** when you don't specify one and covers all four types — **serverless / standard / stateful**, and **cron** by setting **`type: cron`** (which makes `schedule` required). Type is chosen at creation and is **immutable** (see Immutability below). Pick from:
| | **standard** (default) | serverless | stateful | cron |
|---|---|---|---|---|
| Use for | long-running services, APIs, workers | request/event-driven HTTP that scales on demand | databases & anything needing stable disk or per-replica identity | scheduled jobs |
| Autoscaling metrics | cpu, memory, latency, rps, multi, keda, disabled | concurrency, cpu, memory, rps, disabled | cpu, memory, latency, rps, multi, keda, disabled | n/a — runs on a `schedule` |
| Capacity AI | on by default | on by default | not applied | not applied |
| Probes | define readiness + liveness | define readiness + liveness | define readiness + liveness | ignored |
| `ext4`/`xfs` volumes | no | no | **yes (only here)** | no |
| `shared` volumes | yes | yes | yes | yes |
| Scale to zero | KEDA only | yes | KEDA only | n/a |
| Default `minScale` | 1 | 1 | 1 | n/a |
The intended scaling metric can decide the type: `concurrency` scaling exists **only on serverless** — if that's the intent, create the workload as serverless (type is immutable); on standard/stateful the closest equivalent is `rps`. Never pair a metric with a type that rejects it.
A workload has **1–8 containers**.
## The spec at a glance — which tool sets what
There is ONE way to express each concept. Containers always go in the typed `containers[]` array (there are no flat `image`/`cpu`/`port` fields), scaling always goes in the single `autoscaling` block, and cron is **`create_workload` / `update_workload` with `type: cron`** — the `schedule` + job policy become available (and required), while autoscaling/`capacityAI`/`timeoutSeconds`/`debug` do not apply to cron and are rejected. The advanced blocks below were split into dedicated `configure_workload_*` tools to keep the common path lean.
| Spec block | What it controls | Set with |
|---|---|---|
| `containers[]` — `image`, `ports`, `cpu`/`memory`, `env`, `command`/`args`, probes, `metrics`, `volumes` | the container(s) — the only way to define them | `create_workload` / `update_workload` (all types, cron included) |
| `autoscaling` (→ `spec.defaultOptions.autoscaling`) + `capacityAI` / `timeoutSeconds` / `suspend` / `debug` scalars | scaling & resource optimization | `create_workload` / `update_workload` |
| `firewallConfig` (or the `public` shortcut) | inbound/outbound/internal exposure | `create_workload` / `update_workload` (all types, cron included) |
| `schedule` + cron policy (`concurrencyPolicy`, `historyLimit`, `restartPolicy`, `activeDeadlineSeconds`) | cron schedule & job policy | `create_workload` / `update_workload` **with `type: cron`** |
| `loadBalancer` (direct / geo / replicaDirect) | custom ports, static IPs, geo headers | `configure_workload_load_balancer` |
| `sidecar.envoy` | Envoy filter chain (e.g. JWT auth) | `configure_workload_sidecar` |
| `extras` | BYOK-only affinity / tolerations / topology | `configure_workload_extras` |
| `localOptions` (incl. `spot`, `multiZone`, `capacityAIUpdateMinutes`) | per-location overrides of `defaultOptions` | `configure_workload_local_options` |
| `rolloutOptions` | graceful termination, surge/unavailable | `configure_workload_rollout` |
| `securityOptions` | `runAsUser`, `filesystemGroupId` | `configure_workload_security` |
| `requestRetryPolicy` | request retry attempts / conditions | `configure_workload_retry` |
`update_workload` merges `containers[]` **by name** — send only the container(s) you want to change; others are preserved (an unknown name adds a container). On a cron workload, `update_workload` patches the `schedule` / job policy / `suspend` / containers (and rejects autoscaling/`capacityAI`/`timeoutSeconds`/`debug`); schedule/job fields are rejected against a non-cron workload. Always call `get_resource_schema` for the workload kind before authoring a spec — never hand-write fields from memory.
## Production-grade defaults
Platform defaults are not a production design. For any real workload:
- **`minScale ≥ 2`** for user-facing services (HA — no single point of failure). The schema default is `1`; use `1` only with a named reason (single-writer DB, leader election, dev/staging). `stateful` is often correct at `1`.
- **`maxScale`**: leave it at the default of **5** unless the user gives an explicit maximum. If the user says "max 10 replicas" (or names any number), set exactly that. Do not invent a different cap.
- **Never set `minScale: 0` (scale-to-zero)** unless the user asks for it by name. `serverless` scales to zero directly; `standard`/`stateful` only with `metric: keda`; `cron` cannot.
- **Define both `readinessProbe` and `livenessProbe`** — none are configured by default.
- **Size `cpu`/`memory` to the runtime**, not the platform defaults (`50m` / `128Mi`). Floors: CPU ≥ `25m`, memory ≥ `32Mi`. Keep `memory(MiB) / cpu(millicore) ≤ 8` (raise to 32 with the tag `cpln/relaxMemoryToCpuRatio`).
- **Pick an autoscaling metric that fits the traffic shape** (see Autoscaling).
- **Set the firewall to match intended exposure IN THE CREATE CALL** — it is deny-by-default (see Networking). Decide reachability before creating (`public: true` or `firewallConfig`); creating closed and patching the firewall open afterward is a spec error, not a workflow.
- **Never silently downgrade** an incompatible request to `disabled` / `none` / `1` / public — surface the conflict with realistic alternatives and a recommendation.
## Images
- Your org's private registry, in a spec: **`//image/NAME:TAG`** (e.g. `//image/api:v1.0`) — the preferred form.
- **Another Control Plane org's registry: `OTHER-ORG.registry.cpln.io/NAME:TAG`** — this hostname form is valid in a workload spec for cross-org pulls.
- Public images: the **exact string** (`nginx:latest`) — **never** add a `docker.io/` prefix. ECR/GCR/etc. use their full host path.
- The `<your-org>.registry.cpln.io/NAME:TAG` form also resolves, but for your own org prefer `//image/NAME:TAG`; the hostname form is mainly used by `docker login` / `docker push`.
- **All images must be `linux/amd64`** — a wrong-arch image fails with `exec format error`.
- **Private external registries need a pull secret on the GVC** (`spec.pullSecretLinks`); only `docker`, `ecr`, and `gcp` secret types work as pull secrets. Same-org `//image/...` needs none.
- **Build and push:** `cpln image build --name NAME:TAG --remote` builds on Control Plane and pushes for you (no Docker daemon); `--push` builds locally. Over MCP, `build_image` starts a build **from a GitHub/GitLab repo only** — a local folder has no path through MCP and must use the CLI.
- Image **records** over MCP are list/get/delete (`list_resources` / `get_resource` / `delete_resource`, kind="image"). Detail: `image` skill.
## Run real images
Run an actual container image — not an inline/base64/heredoc app on a generic base image. For databases, caches, queues, brokers, search, gateways, or other common infrastructure, install a Template Catalog entry first (`browse_templates` → `install_template`) rather than hand-building.
## Health, readiness & verification
- **`readinessProbe` gates traffic** — the load balancer only routes to a ready replica. It should check request-path dependencies (DB, auth, cache).
- **`livenessProbe` restarts a hung process** — it must check **only** the process itself, never downstream dependencies (a dependency outage must not cycle every replica).
- Each probe is exactly one of `exec` / `grpc` / `tcpSocket` / `httpGet`. Tune `initialDelaySeconds` to real cold-start time (readiness default 10s, liveness default 60s; `periodSeconds` default 10s).
- **Verify every create/update automatically — without asking:** poll `list_deployments` until all locations report ready (it surfaces per-location errors **and** the workload's canonical public URL). Then give the user that **canonical** URL — never construct one or report a per-location deployment URL as the address. **For a public workload, do not stop at "ready" — confirm it actually serves:** make a real HTTP GET of the canonical endpoint (when you have that capability) and read the result — never claim reachability without a real response you received; if you cannot make a request, report readiness confirmed but external reachability not independently verified. A ready deployment can still be unreachable — firewall inbound unset, or TLS/DNS still propagating. Treat 2xx/3xx/401/403 as serving; a timeout/refused points first at firewall inbound, a TLS/DNS error at propagation (wait, don't redeploy). On failure, diagnose with `get_workload_events` (probe/scheduling reasons) then `get_workload_logs` (app error); pass the optional `location` to `list_deployments` (e.g. `aws-us-east-1`) to inspect ONE location's deployment in full detail. **Never re-apply an unchanged failing spec**, and don't poll in a tight loop.
## Autoscaling & capacity
Set via `spec.defaultOptions.autoscaling.metric`; the system keeps the metric near but below `target` (default `95`; capped at 100 for cpu/memory). If `metric` is omitted, serverless defaults to `concurrency` and standard/stateful default to `cpu`. Picker:
- **concurrency** — HTTP with variable request duration (**serverless only**).
- **rps** — HTTP with consistent response times.
- **cpu** / **memory** — compute- or memory-bound work.
- **latency** — SLO-driven APIs (**standard / stateful**; set `metricPercentile`).
- **multi** — several signals, highest replica count wins (**standard / stateful**; entries limited to `cpu`/`memory`/`rps`; mutually exclusive with `metric`/`target`).
- **keda** — event-driven (queues/streams; **standard / stateful**); requires `spec.keda.enabled: true` on the GVC; `target` is rejected with `keda`.
- **disabled** — fixed replicas at `minScale`.
The metric must be valid for the workload type (the matrix above) or the spec is rejected — e.g. `concurrency` on a `standard` workload is rejected (it is serverless-only). Match the metric to the workload's traffic shape: `rps`/`concurrency` for HTTP, `cpu`/`memory` for compute-bound work, `latency` for SLO-driven APIs. For tuning targets/percentiles, multi-metric, KEDA, scale-to-zero, or Capacity AI, load the `autoscaling-capacity` skill.
**Capacity AI** auto-tunes CPU/memory between `minCpu`/`minMemory` and `cpu`/`memory`. On by default for **standard** and **serverless**; **not applied** to stateful or cron. It is **rejected with the `cpu` metric** (when explicitly enabled) and **with GPUs**.
## Networking, firewall & exposure
- **Deny-by-default:** external inbound, external outbound, and internal (`inboundAllowType: none`) are all blocked until configured. Blocked CIDRs beat allowed; CIDR rules beat hostname rules.
- **Public exposure needs BOTH** an external inbound and an external outbound CIDR — one without the other ships a half-broken workload. Infer intent: a user-facing app/site/game → public; an internal API/DB/worker → restricted. Confirm when ambiguous or sensitive — and decide BEFORE creating: exposure belongs in the create call itself, never a follow-up firewall patch.
- Hostname outbound rules allow only ports **80/443/445** by default; `outboundAllowPort` **replaces** that set (re-list 80/443 if still needed). Private RFC1918/CGNAT ranges in `outboundAllowCIDR` are silently ignored on managed locations — reaching private networks takes a wormhole agent (`native-networking`).
- **Internal service-to-service** uses plain HTTP over the internal hostname: `http://WORKLOAD.GVC.cpln.local:PORT` (the sidecar adds mTLS — never `https://`). Same-GVC is free; cross-GVC needs `inboundAllowType: same-org` (or an explicit `workload-list`) and incurs egress.
- **One public canonical port.** `WORKLOAD.GVC.cpln.app` serves a **single** port — the first container port. `standard`/`stateful` may expose **more** ports across containers (unique numbers), reachable at `WORKLOAD.GVC.cpln.local:PORT` or via a **direct/dedicated load balancer**; `serverless` is limited to one container / one port. `WORKLOAD.GVC.cpln.app` is the URL *shape* only — always report the **actual** canonical URL from `list_deployments` / the workload's `status.canonicalEndpoint`; never construct or guess it (custom domains, BYOK, and alias suffixes make the literal form wrong).
- **Always declare ports with the `containers[].ports` array** — e.g. `ports: [{ number: 80, protocol: "http" }]`; for a single port use a one-element array. The legacy scalar `containers[].port` field is **deprecated — never use it**, even if `get_resource_schema` still lists it (the platform keeps it for backward compatibility, but new specs must use `ports[]`).
- **Load balancer picker:** shared (default, HTTP/HTTPS on 80/443, no config) · **direct** — per-workload custom TCP/UDP `externalPort` 22–32768, optional static IPs via an IP set, geo headers; set with `configure_workload_load_balancer` · **dedicated** — per-GVC custom domains and wildcard hosts; a GVC setting, enabled with `update_gvc`. `firewallConfig` stays on `create_workload` / `update_workload`. Toggling direct/dedicated needs the `configureLoadBalancer` permission — `edit` does not imply it (`ipset-load-balancing` skill).
## Persistent storage
- Need durable disk or stable per-replica identity → a **`stateful`** workload with a mounted **volume set**.
- A volume set's **filesystem** (`ext4` / `xfs` / `shared`) and **performance class** are **immutable** — set at creation.
- `ext4`/`xfs` mount on **stateful only**; `shared` mounts on any type. Up to **15 volumes per container**, and no two mounts in a container may share a path or nest (one mount path cannot be a parent of another). Reserved mount paths (rejected): `/dev`, `/dev/log`, `/tmp`, `/var`, `/var/log`.
- **Snapshot before any destructive volume op** (shrink/restore/delete); snapshots exist for `ext4`/`xfs` only.
- Attach with `mount_volumeset_to_workload`.
## Secrets, env vars & naming rules
- **Secrets** can be consumed two ways: as an **environment variable value** — `cpln://secret/NAME` (or `cpln://secret/NAME.key` for a keyed/dictionary secret) — or **mounted as a volume** with `uri: cpln://secret/NAME`. Either way the workload still needs **all three pieces**: an identity on the workload, a policy granting `reveal`, and the reference — or access fails silently. `grant_workload_secret_access` sets the identity + policy but not the reference, and it requires the workload to **already exist** — for a new workload, `create_workload` first (its deployment pauses on the secret reference until access is granted, then resumes).
- **Environment variable names cannot start with `CPLN_`** (reserved). The platform injects these at runtime: `CPLN_TOKEN`, `CPLN_ENDPOINT`, `CPLN_GLOBAL_ENDPOINT`, `CPLN_ORG`, `CPLN_GVC`, `CPLN_GVC_ALIAS`, `CPLN_LOCATION`, `CPLN_PROVIDER`, `CPLN_WORKLOAD`, `CPLN_WORKLOAD_VERSION`, `CPLN_IMAGE`, `CPLN_NAME` (plus `CPLN_MAIN` on the first container, and `PORT` on standard when unset). `K_SERVICE` / `K_CONFIGURATION` / `K_REVISION` are also disallowed. Names match `^[-._a-zA-Z][-._a-zA-Z0-9]*$` (max 120 chars).
- **A workload can call the Control Plane API as its identity:** `curl -H "Authorization: Bearer $CPLN_TOKEN" $CPLN_ENDPOINT/org/$CPLN_ORG/...` — `CPLN_ENDPOINT` is plain **http** (the sidecar secures and signs it in transit). Requests act as the attached `spec.identityLink` identity and succeed only where a policy grants that identity the permission — no identity attached or no policy means 403. The token works **only from inside that workload, against `CPLN_ENDPOINT`**: it does not authenticate to `api.cpln.io`, `metrics.cpln.io`, or `logs.cpln.io` (use a service-account key there).
- **Container names cannot start with `cpln-` or `debugger-`** (and a few exact names like `istio-proxy` are reserved). Names are lowercase `^[a-z]([-a-z0-9])*[a-z0-9]$`, max 64.
- **Workload name** is max **49** characters, cannot end with `-headless`, and is immutable.
## Runtime traps
- **Graceful shutdown:** the default `preStop` runs `sh -c "sleep N"`. Minimal/distroless images often lack `sleep` — if it (or a custom `preStop`) fails in **any** container, **all** containers are SIGKILL'd immediately. Grace period is `spec.rolloutOptions.terminationGracePeriodSeconds` (0–900, default 90).
- **Reserved container ports** (rejected): `8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000`. Valid container port range is 80–65535; **port numbers must be unique across all containers**. A **serverless** workload must expose **exactly one port, on exactly one container**. Declare every port in the `containers[].ports` array (`[{ number, protocol }]`) — the scalar `containers[].port` field is deprecated; do not use it.
- **Don't run as UID 1337** — that is the mesh proxy's UID. A container with `runAsUser: 1337` has its outbound traffic excluded from the Envoy sidecar redirect, so it bypasses the mesh — losing mTLS and firewall enforcement (it gets *unfiltered* egress, not "no networking").
## Immutability & destructive changes
- **Workload `type` and `name` are immutable.** Changing either = **delete + recreate**, which is **destructive**: it drops the public URL `WORKLOAD.GVC.cpln.app`, the internal DNS `WORKLOAD.GVC.cpln.local`, and policy `targetLinks` / identity bindings. Recreating with the **same name** preserves the URL/DNS; a different name silently breaks every external reference.
- The same applies to a volume set's filesystem and performance class.
- Before any delete or immutable-forcing change, present **Action · Affected · Blast radius · Data / Traffic / Access impact · Reversibility · Mitigation** and wait for explicit confirmation (see the root rules).
## Metrics & observability
- Built-in metrics (CPU/memory reserved-vs-used, request rate/latency, replica count, restarts) exist for every workload with no config.
- Custom Prometheus: add `spec.containers[].metrics` with `port` (required) and `path` (default `/metrics`).
- Query with `list_metrics` (discover real names/labels) → `query_metrics` (PromQL). Confirm a signal exists before changing scaling.
## Running commands in a live replica
Use `list_workload_replicas` for replica discovery. When in-container inspection is essential, read the `cpln` skill and use its verified CLI workflow. Any state-changing command needs explicit user confirmation after stating the exact command, impact, and risk; never surface resolved secret values.
## Standard create / update flow
1. Read this skill once per session before authoring (you are doing that now); `get_cpln_rules` has the cross-cutting operating guide if you have not read it this session.
2. Confirm the target **org / GVC** — never guess; on not-found, stop and ask.
3. `get_resource_schema` for the workload kind before authoring.
4. Discover current state: `list_resources` (kind="workload") / `get_resource` (kind="workload").
5. Prepare the smallest valid change; if destructive, confirm.
6. `create_workload` / `update_workload` (for a scheduled job, pass `type: cron` with a `schedule`; PATCH — only sent fields change, containers merged by name), plus `configure_workload_*` for load balancer / sidecar / extras / local options / rollout / security / retry.
7. Verify automatically — do not ask permission: poll `list_deployments` until every location is ready; on failure diagnose with events → logs and fix.
8. Report exactly what changed and the resulting status — and for an exposed workload, give the user its **canonical** public URL (read from `list_deployments` or the workload's `status.canonicalEndpoint`; never construct/guess it or report a per-location URL as the address).
## Quick reference — MCP tools
| Tool | Purpose |
|---|---|
| `create_workload` | Create any workload (typed `containers[]`, single `autoscaling` block) — including a scheduled job with `type: cron` + a required `schedule`. |
| `update_workload` | Update a workload (PATCH; containers merged by name) — on a cron workload, patches `schedule` / job policy / `suspend`. |
| `get_resource` (kind="workload") / `list_resources` (kind="workload") | Read one / list in a GVC (capture state before changes). |
| `delete_resource` (kind="workload") | Delete a workload (destructive — confirm blast radius first). |
| `configure_workload_load_balancer` | Set/clear `spec.loadBalancer` (direct, geo headers, replicaDirect). |
| `configure_workload_sidecar` | Set/clear `spec.sidecar.envoy` (Envoy filters, JWT auth). |
| `configure_workload_extras` | Set/clear `spec.extras` (BYOK affinity/tolerations/topology). |
| `configure_workload_local_options` | Set/clear `spec.localOptions` (per-location overrides). |
| `configure_workload_rollout` | Set/clear `spec.rolloutOptions` (graceful termination, surge/unavailable). |
| `configure_workload_security` | Set/clear `spec.securityOptions` (`runAsUser`, `filesystemGroupId`). |
| `configure_workload_retry` | Set/clear `spec.requestRetryPolicy` (retry attempts/conditions). |
| `list_deployments` | PRIMARY post-deploy readiness monitor (all locations); per-location errors **and** the canonical public URL to report. Pass the optional `location` (e.g. `aws-us-east-1`) for ONE deployment's full detail — version chain, per-container readiness, full JSON. |
| `get_workload_events` | Probe/scheduling failures after a bad deploy. |
| `get_workload_logs` | App-side logs (LogQL) for runtime/startup errors. |
| `list_workload_replicas` | List running replicas. |
| `workload_start_cron` | Trigger an out-of-band run of a cron workload. |
| `grant_workload_secret_access` | Grant an **existing** workload secret access (identity + `reveal` policy; you still add the reference — create the workload first). |
| `mount_volumeset_to_workload` | Attach a volume set to a stateful workload. |
**CLI fallback** (read the `cpln` skill first): use when MCP is unavailable/unauthenticated, for live container commands and interactive work, a local-folder image build or an image copy, or as the primary interface in CI/CD (`CPLN_TOKEN` + `cpln apply --ready`).
**Raw API escape hatch:** for a spec field no typed `create_workload` / `update_workload` / `configure_workload_*` tool exposes, use `cpln_api_request` (raw GET/POST/PATCH/DELETE; disabled by default — only when advertised) — call `get_resource_schema` first for the exact path and body, and prefer the typed tools whenever they cover the field. If it is not advertised, apply the full manifest with the `cpln` CLI instead.
## Deep-dive router
Load the matching skill (one or several) when you need more than the primary rules above:
| Need | Skill |
|---|---|
| Image refs, builds, buildpacks, registries, pull secrets, cross-org sharing | `image` |
| Autoscaling, Capacity AI, scale-to-zero, KEDA, custom-metric scaling | `autoscaling-capacity` |
| Probes in depth, JWT/Envoy auth, security options, graceful termination | `workload-security` |
| Firewall rules, inbound/outbound, header & geo filtering | `firewall-networking` |
| Static IPs, direct & dedicated load balancers, custom ports | `ipset-load-balancing` |
| CDN caching, request rate limiting, DDoS protection | `cdn-rate-limiting` |
| Volumes, volume sets, snapshots, persistence, expansion | `stateful-storage` |
| Metrics, PromQL, Grafana, Prometheus federation | `metrics-observability` |
| Logs, LogQL, events, per-execution cron logs | `logql-observability` |
| Private networking, agents, VPC, on-prem connectivity | `native-networking` |
| Running workloads on own hardware, bare metal, data center, mk8s, BYOK | `mk8s-byok` |
| Databases, caches, queues, brokers, common infra | `template-catalog` |
| Secrets, identities, policies, RBAC, service accounts | `access-control` |
## Documentation
- [Workload Reference](https://docs.controlplane.com/reference/workload/general.md)
workload-security13.3 KB
---
name: workload-security
description: "Production hardening for Control Plane workloads. Use when asked about JWT/Envoy auth, security context (runAsUser), health probe tuning, direct load balancers, geo-location headers, or graceful shutdown / termination."
---
# Workload Security & Production Hardening
Deep-dive companion to the `workload` skill, which owns workload types, the spec shape, and the readiness-vs-liveness model. Everything below is production-hardening detail for an existing workload.
**Where settings live.** Health probes go inline in `containers[]` via `create_workload` / `update_workload`. Every other block here — `sidecar.envoy`, `loadBalancer`, `securityOptions`, `rolloutOptions` — is set by its own `configure_workload_*` tool, a set-or-clear PATCH on that one field (`remove: true` clears it).
## Health Probes
Define `readinessProbe` (gate traffic) and `livenessProbe` (restart on failure) as distinct probes in the container spec.
**Defaults by workload type:**
- **Serverless** — a TCP readiness probe on the listening port is injected by default (plus default startup and liveness probes). Adequate, but an `httpGet` against a real endpoint catches more failure modes (DB unreachable, dependency timeout, deadlock).
- **Standard / Stateful** — **no probes by default**; add them explicitly for any production workload.
- **Cron** — probes are stripped (ignored).
No HTTP healthcheck? Use `tcpSocket` on the listening port as a baseline — don't run a long-lived workload probe-less.
### Probe schema
Each probe takes exactly one of `exec` / `grpc` / `tcpSocket` / `httpGet`, plus these timing fields:
| Field | Range | Default |
|---|---|---|
| `initialDelaySeconds` | 0-600 | 10 (readiness) / 60 (liveness) |
| `periodSeconds` | 1-600 | 10 |
| `timeoutSeconds` | 1-600 | 1 |
| `successThreshold` | 1-20 | 1 |
| `failureThreshold` | 1-20 | 3 |
`httpGet` omitting `port` defaults to the first container port; `httpGet.scheme` defaults to `HTTP`. Keep liveness looser than readiness (e.g. `periodSeconds: 30`) — restarts are expensive.
```yaml
containers:
- name: api
image: //image/api:v1.0
ports: [{ number: 8080, protocol: http }]
readinessProbe:
httpGet: { path: /healthz/ready, port: 8080 }
initialDelaySeconds: 5
failureThreshold: 3
livenessProbe:
httpGet: { path: /healthz/live, port: 8080 }
initialDelaySeconds: 30
periodSeconds: 30
```
## JWT Authentication
JWTs are validated at the Envoy sidecar before requests reach the workload, via the `jwt_authn` HTTP filter under `spec.sidecar.envoy`. Set it per-workload with `configure_workload_sidecar`, or org-wide-per-GVC by putting the same `sidecar.envoy` on the GVC (`update_gvc`) — it then applies to every workload in that GVC.
Must-know rules (the filter is strictly validated, not passthrough):
- The filter `name`, `typed_config."@type"`, and `priority` (0-100) must be exact — copy them from the example.
- Each provider needs a matching `clusters[]` entry (`STRICT_DNS` + TLS transport socket) so Envoy can fetch the JWKS over HTTPS. `remote_jwks.http_uri.cluster` must equal that cluster's `name`.
- `rules` are first-match-wins. A rule with no `requires` lets matching paths through **without** a token — use it to exempt health/metrics endpoints.
- `claim_to_headers` forwards JWT claims into request headers, so the workload gets identity context without re-parsing the token.
- Provider/cluster names starting with `cpln_` are UI-managed and restricted (no `async_fetch` / `retry_policy`, and `cache_duration` must equal `http_uri.timeout`). For hand-authored configs use a **non-`cpln_`** name for full Envoy flexibility.
```yaml
spec:
sidecar:
envoy:
clusters:
- name: auth0
type: STRICT_DNS
load_assignment:
cluster_name: auth0
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: YOUR_TENANT.auth0.com, port_value: 443 }
transport_socket:
name: envoy.transport_sockets.tls
http:
- name: envoy.filters.http.jwt_authn
priority: 50
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.jwt_authn.v3.JwtAuthentication
providers:
auth0:
issuer: https://YOUR_TENANT.auth0.com/
audiences: [https://api.example.com]
remote_jwks:
http_uri: { uri: https://YOUR_TENANT.auth0.com/.well-known/jwks.json, cluster: auth0, timeout: 5s }
cache_duration: 300s
claim_to_headers:
- { header_name: X-User-Sub, claim_name: sub }
rules:
- match: { prefix: /healthz } # public, no token required
- match: { prefix: / }
requires: { provider_name: auth0 } # everything else needs a valid JWT
```
For any OIDC provider (Auth0, Firebase, Cognito, Okta), set `issuer` to its issuer URL, `remote_jwks.http_uri.uri` to its JWKS endpoint (usually `/.well-known/jwks.json`), and `audiences` to your client ID or API identifier.
## Security Options
`spec.securityOptions` (set with `configure_workload_security`):
| Field | Range | Purpose |
|---|---|---|
| `runAsUser` | 1-65534 | UID for all container processes |
| `filesystemGroupId` | 1-65534 | GID applied to mounted volumes |
Neither has a schema default — unset means the image's user and root/GID 0 for volumes. Set `runAsUser` to a non-root UID for defense in depth; set `filesystemGroupId` so containers can share access to mounted volumes (e.g. volume sets). **Not valid for `type: vm`** (the guest OS owns its own security context).
**Trap — never `runAsUser: 1337`.** That is the mesh proxy's UID. A container running as 1337 is excluded from the Envoy redirect, so it bypasses the mesh entirely — losing mTLS and firewall enforcement and getting *unfiltered* egress.
### Workload permissions
Beyond standard `view` / `edit` (implies `view`) / `create` / `delete` / `manage`, the workload kind adds three non-obvious permissions: `connect` (interactive shell into a replica) and `configureLoadBalancer` (toggle direct/dedicated LBs) — **neither implied by `edit`** — plus `exec`, which implies `exec.runCronWorkload` + `exec.stopReplica`. Full policy and principal setup: `access-control` skill.
## Direct Load Balancers
`spec.loadBalancer.direct` exposes workload ports through a per-location cloud LB. Unlike the shared LB, it supports custom TCP/UDP ports and needs no domain registration — use it for non-HTTP protocols or custom ports. You are responsible for your own TLS certificates. Set with `configure_workload_load_balancer`.
- `enabled` (bool, required) — when `false`, the LB is stopped and accrues no charges.
- `ports[]` — each entry: `externalPort` (22-32768, required), `protocol` (`TCP`/`UDP`, required), `containerPort` (80-65535, excludes reserved ports), `scheme` (`http`/`tcp`/`https`/`ws`/`wss`, optional, sets the UI link scheme; default `https`).
- `ipSet` (optional) — link to an IP set for reserved static IPs.
**Reserved container ports (rejected):** 8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000.
```yaml
spec:
loadBalancer:
direct:
enabled: true
ports:
- { externalPort: 5432, protocol: TCP, scheme: tcp, containerPort: 5432 }
```
Each direct-LB location also gets a public Geo DNS address with latency-based routing, usable as a CNAME target for custom domains.
### Geo location headers
`spec.loadBalancer.geoLocation` adds MaxMind GeoLite2 data to inbound HTTP requests under the header names you set in `headers.{asn,city,country,region}` (each max 128 chars; `enabled` defaults `false`). When enabled, at least one header is required, names must be unique, and existing same-named headers are replaced. HTTP-exposed workloads only. Pair with header-based firewall rules for geo-filtering (`firewall-networking` skill).
### Replica-direct endpoints
`spec.loadBalancer.replicaDirect: true` gives each replica its own stable hostname (default `false`, **stateful workloads only**):
- External: `WORKLOAD-GVCALIAS-INDEX.LOCATION.controlplane.us`
- Internal: `replica-INDEX.WORKLOAD.LOCATION.GVC.cpln.local:PORT`
## Graceful Termination
`spec.rolloutOptions` (set with `configure_workload_rollout`) governs how replicas are removed during scaling, version updates, Capacity AI rollouts, and maintenance.
| Field | Range / values | Default |
|---|---|---|
| `minReadySeconds` | integer ≥0 | 0 |
| `maxSurgeReplicas` | integer or percent | unset (not for `vm`) |
| `maxUnavailableReplicas` | integer or percent | unset (not for `stateful`) |
| `scalingPolicy` | `OrderedReady` / `Parallel` | `OrderedReady` (not for `vm`) |
| `terminationGracePeriodSeconds` | 0-900 (3600 with the `cpln/relaxGracePeriodMax` tag) | 90 |
`terminationGracePeriodSeconds` is the total budget for a replica to shut down before SIGKILL; raise it for long-running requests. Termination sequence:
1. The load balancer removes the replica from the pool (up to ~10s) so new requests route elsewhere.
2. The Control Plane sidecar drains in parallel: it holds for `grace - 10` seconds (default 80), then waits for in-flight connections to complete before shutting down.
3. Containers run their `preStop` hook. With no custom hook, a default `sh -c "sleep N"` runs where `N` = half the grace period (45s at the 90s default), giving the LB time to stop routing.
4. After `preStop`, the container receives **SIGINT** and has the remaining grace to exit, then **SIGKILL**. Handle the termination signal to drain cleanly.
**Trap — immediate SIGKILL of *all* containers** if `sleep`/`sh` is missing in *any* container (common with distroless/minimal images) or a custom `preStop` errors in *any* container. A custom `preStop` must include a delay or connection check, and must be tested — an error there force-kills the whole replica.
## Verify
- After any change, poll `list_deployments` until each location reports ready — it surfaces probe failures per location.
- `get_workload_events` gives the probe/liveness failure reason and message; `get_workload_logs` shows app-side errors.
- For JWT, send a request with and without a valid token (expect 401 without) and confirm the `claim_to_headers` header arrives at the workload.
## Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Deploy never ready; events show probe failures | wrong `httpGet` path/port, or app slow to start | fix path/port; raise `initialDelaySeconds` / `failureThreshold` |
| All requests 401 after adding JWT | no public-exemption rule, or `provider_name` mismatch | add a no-`requires` rule for health paths; match the rule's `provider_name` to a provider key |
| JWT always rejected / JWKS fetch fails | cluster missing or `remote_jwks.http_uri.cluster` doesn't match a `clusters[].name` | add the `STRICT_DNS` + TLS cluster and align the names |
| Replica SIGKILL'd instantly on rollout | `sleep`/`sh` missing, or custom `preStop` errors | use a `sleep`-capable image, or a native-sleep `preStop`; test the hook |
| Direct LB port rejected | `containerPort` is reserved, or `externalPort` outside 22-32768 | pick a non-reserved container port; keep `externalPort` in range |
| `securityOptions` / `replicaDirect` rejected | `type: vm` (securityOptions) or non-stateful (replicaDirect) | remove the field or change the workload type |
## Quick reference
### MCP tools
All `configure_workload_*` tools are `full`-profile, set-or-clear PATCH (`remove: true` clears).
| Tool | Purpose |
|---|---|
| `configure_workload_sidecar` | Set/clear `spec.sidecar.envoy` — JWT / Envoy filter chain |
| `configure_workload_load_balancer` | Set/clear `spec.loadBalancer` — direct LB, geo headers, replicaDirect |
| `configure_workload_security` | Set/clear `spec.securityOptions` — `runAsUser`, `filesystemGroupId` |
| `configure_workload_rollout` | Set/clear `spec.rolloutOptions` — termination grace, surge/unavailable |
| `create_workload` / `update_workload` | Probes go inline in `containers[]` (update is PATCH, merges containers by name) |
| `list_deployments` | Poll per-location readiness; surfaces probe failures |
| `get_workload_events` | Probe / liveness failure reason + message |
| `get_workload_logs` | App-side logs for security / probe issues |
### CLI (fallback)
Use the CLI when the MCP server is unavailable or unauthenticated, or in CI/CD (service-account `CPLN_TOKEN`).
```bash
cpln workload get WORKLOAD --gvc GVC -o yaml-slim > workload.yaml
# edit, then apply (CI/CD: add --ready to block until deployed)
cpln apply -f workload.yaml --gvc GVC
```
### Related skills
- `workload` — primary skill (types, defaults, spec shape, tool division); start here.
- `firewall-networking` — CIDR / header rules, geo-filtering, LB types.
- `access-control` — policies, principals, the full permission model.
- `autoscaling-capacity` — scaling and Capacity AI settings.
- `stateful-storage` — volume sets (pair with `filesystemGroupId`).
## Documentation
- [Workload Security](https://docs.controlplane.com/reference/workload/security.md)
- [JWT Auth](https://docs.controlplane.com/reference/workload/jwt-auth.md)
- [Termination](https://docs.controlplane.com/reference/workload/termination.md)
- [Load Balancing](https://docs.controlplane.com/reference/workload/load-balancing.md)
workload-troubleshooting15.8 KB
---
name: workload-troubleshooting
description: "Diagnoses unhealthy Control Plane workloads. Use when asked why a workload is crashing, not starting, OOMKilled, ImagePullBackOff, returning 502s, failing health checks, unreachable, or stuck deploying."
---
# Workload Troubleshooting
The symptom-first companion to the `workload` skill (which owns workload types, the spec shape, and the create/update tools). Given an unhealthy workload, map what you observe to its platform-specific root cause and a fix the schema will actually accept. Diagnosis is **read-only and MCP-first**; most failures trace to a Control Plane rule a generic engineer would not guess — deny-by-default firewalls, the secret identity+policy chain, blocked ports, the sleep-binary shutdown rule. The single most common is OOMKilled. Deep remediation for each area lives in the domain skill named in that section; this skill is the diagnostic map.
## Step 1 — Gather state (read-only)
| Tool | What it tells you |
|---|---|
| `list_deployments` | **Start here.** Per-location readiness with reason/message. Pass `location` to drill into one failing location. |
| `get_workload_events` | Image pulls, crashes, scheduling, probe failures, `OOMKilled`. |
| `get_workload_logs` | App logs (LogQL); the `_accesslog` container holds HTTP status codes and latency. |
| `get_resource` (kind=`workload`) | The spec and current status. |
| `list_metrics` then `query_metrics` | Resource pressure — memory before OOM, CPU, latency. |
| `list_workload_replicas` | Confirm which replicas are currently running. Use the `cpln` CLI after reading the `cpln` skill if in-container inspection is essential. |
| `query_traces` then `get_trace` | For a slow or intermittently failing request: which span in the path spent the time or errored. Needs tracing enabled on the GVC (opt-in). Deep dive in `metrics-observability`. |
CLI fallback (MCP unavailable, interactive shell, or CI/CD):
```bash
cpln workload get WORKLOAD --gvc GVC -o json
cpln workload eventlog WORKLOAD --gvc GVC -o json
cpln logs '{gvc="GVC", workload="WORKLOAD"}' --limit 50 # |= "error" filters; container="_accesslog" for HTTP codes
cpln workload connect WORKLOAD --gvc GVC --location LOCATION # interactive shell
```
## Failure catalog
### Out of memory (OOMKilled) — the most common issue
**Symptoms:** container restarts repeatedly, events show `OOMKilled`, crashes under load.
`memory` is a hard cap — exceed it (app + runtime + GC + buffers) and the kernel kills the container. Usual culprits: Java without `-Xmx`, Node without `--max-old-space-size`, Python loading large datasets. Each container, sidecars included, has its own limit. With **Capacity AI** on, spiky workloads can be downsized too aggressively — set `minMemory` as a floor (Capacity AI never downscales CPU below 25 millicores).
**Fix:** check real usage with `query_metrics`, then raise `memory` — but memory (MiB) must stay within **8× CPU (millicores)**, so raise `cpu` alongside it or the update is rejected (the default `cpu: 50m` caps memory at 400Mi). See *Apply fixes within the schema's limits*.
### Image pull failures
**Symptoms:** events show `ImagePullBackOff` / `ErrImagePull`, deployment stuck.
- **Reference format** — `//image/NAME:TAG` (org registry), bare `NAME:TAG` (Docker Hub, no `docker.io/`), full URL (other registries).
- **Platform** — images must be `linux/amd64` for managed locations (BYOK allows more).
- **Pull secret** — a private external registry needs a pull secret in the GVC's `pullSecretLinks`; only `docker`, `ecr`, `gcp` secret types work, and org `//image/` images need none.
The registry secret must already exist (created by the user — offer a manifest scaffold for them to fill and apply, `setup-secret` skill); attach with `update_gvc`. Deep setup: `image` skill.
### Secret access failures
**Symptoms:** env vars empty, logs show missing config, secret-access errors in events.
A workload reaches a secret only with all three in place: an **identity linked** to it (`spec.identityLink`), a **policy granting that identity `reveal`** on the secret, and a **correct reference**. Fastest fix: `grant_workload_secret_access` builds the whole chain in one call. Reference format by type:
| Type | Reference |
|---|---|
| Opaque (decoded / raw) | `cpln://secret/NAME.payload` / `cpln://secret/NAME` |
| Dictionary | `cpln://secret/NAME.KEY` |
| Username & password | `cpln://secret/NAME.username` / `.password` |
| Keypair | `cpln://secret/NAME.secretKey` / `.publicKey` / `.passphrase` |
| TLS | `cpln://secret/NAME.cert` / `.key` / `.chain` |
| AWS | `cpln://secret/NAME.accessKey` / `.secretKey` / `.roleArn` / `.externalId` |
The manual chain: `access-control` and `setup-secret` skills.
### Port mismatch — healthy but 502/503
- The spec port must match what the process listens on — compare the workload spec with application startup logs. If that is inconclusive, use the `cpln` CLI after reading the `cpln` skill for an in-container socket check.
- On serverless, the runtime injects `PORT` and rejects a `PORT` env var that doesn't equal the exposed port.
- **Type rules:** serverless exposes exactly one port (zero is rejected; TCP needs a dedicated LB — see *Dedicated load balancer & domain*); standard and stateful expose zero or more; cron serves no endpoint.
- **Blocked ports** (cannot bind, invalid for TCP probes): `8012, 8022, 9090, 9091, 15000, 15001, 15006, 15020, 15021, 15090, 41000`.
### Firewall blocking traffic
**Symptoms:** unreachable externally, can't reach external APIs, or can't talk to other workloads.
Deny-by-default: external inbound disabled, external outbound disabled, internal `none`. Fix via `update_workload`:
- **Inbound** — `external.inboundAllowCIDR` (e.g. `0.0.0.0/0`, or specific CIDRs).
- **Outbound** — `external.outboundAllowCIDR`, or `outboundAllowHostname` (hostname rules reach only ports 80, 443, 445).
- **Internal** — `internal.inboundAllowType`: `same-gvc` / `same-org` / `workload-list` (default `none` blocks all workload-to-workload traffic).
Full model: `firewall-networking` skill.
### Health-check failures
**Symptoms:** events show probe failures, replicas unready, restarts.
Default probes: serverless gets readiness + liveness TCP on the container port; standard, stateful, and cron have none (cron strips them). Common fixes: raise `initialDelaySeconds` (0-600; default 10 readiness / 60 liveness) for slow starts; raise `periodSeconds` (1-600, default 10) or `timeoutSeconds` (1-600, default 1) for over-aggressive probes; ensure an HTTP path returns 200-399. **Readiness** failure removes the replica from the pool and pauses rollout; **liveness** failure restarts it. An autoscaled workload with no real readiness probe gets traffic before it's ready, causing 502s on scale-up — add an `httpGet` readiness probe (it needs a port; defaults to the container's). Probe tuning: `workload-security` skill.
### Resource limits & Capacity AI
**Symptoms:** won't schedule, throttled, or Capacity AI not adjusting.
- CPU/memory must fit org quota; `maxScale` × per-replica resources is enforced at scheduling.
- **Capacity AI** does not apply with CPU-utilization or multi-metric autoscaling, or on stateful workloads — use `minCpu` / `minMemory` instead (those persist; `capacityAI` is stripped on stateful).
- **Stateful sizing** — `minCpu`/`cpu` at most 4000m apart (ratio ≤ 4:1); `minMemory`/`memory` at most 4096Mi apart (ratio ≤ 4:1).
- **Ephemeral storage** — 1GB per CPU core (minimum 1GB); exceeding it replaces the replica.
Deep model: `autoscaling-capacity` skill.
### Container won't start (restrictions)
- **UID 1337** is the mesh proxy's UID — running as it excludes the container from the sidecar, disabling mesh communication and mTLS. Override `runAsUser` to another UID in 1-65534 (0/root is rejected).
- **Reserved container names** — not `istio-proxy` / `queue-proxy` / `istio-validation` (or other reserved names); cannot start with `cpln-` or `debugger-`.
- **Reserved env vars** — names starting `CPLN_`, plus `K_SERVICE` / `K_CONFIGURATION` / `K_REVISION`, are rejected; each value caps at 4096 characters.
- **Suspended** — `spec.defaultOptions.suspend: true` stops the workload (min/max scale 0). Clear it, or `cpln workload start WORKLOAD --gvc GVC`.
### Autoscaling misconfiguration
**Symptoms:** won't scale, 502s on scale-up, scale-to-zero not working, or an invalid-strategy error on create.
Strategies: `concurrency`, `cpu`, `memory`, `rps`, `latency`, `keda`, `disabled`. Per type: **serverless** has no `latency` or multi-metric; **standard** has no `concurrency`; **stateful** has no `concurrency`; **cron** has no autoscaling at all (the block is removed). **Scale-to-zero:** serverless with `rps` or `concurrency`; standard and stateful only with `metric: keda` (otherwise the update is rejected); cron cannot. For 502s on scale-up, fix the readiness probe (above). Details: `autoscaling-capacity` skill.
### Termination / graceful shutdown
**Symptoms:** requests fail during deploys or scale-down, 502/503 on rollout, containers killed abruptly.
- **Missing `sleep`** — if `sleep` is absent from **any** container, **all** containers get SIGKILL immediately (no drain). Many distroless/minimal images lack it; confirm with `which sleep`. Fix: include `sleep` or add a custom `preStop`.
- **preStop error** — a failing custom `preStop` in any container SIGKILLs all of them.
- **Ignores SIGTERM** — after the preStop (default `sleep 45`) the container gets the termination signal, then SIGKILL once `terminationGracePeriodSeconds` (default 90; max 900 without the `cpln/relaxGracePeriodMax` tag) elapses.
Sequence and rollout options: `workload-security` skill.
### Volume mount failures
**Symptoms:** can't read mounted files, permission denied, empty cloud volume.
- **Secret volumes** need the identity + `reveal` policy chain (as *Secret access failures*).
- **Cloud volumes** (S3, GCS, Azure Blob/Files) need an identity, a cloud-access policy, and outbound firewall to the provider hosts (`*.amazonaws.com`, `*.googleapis.com`, `*.blob.core.windows.net` / `*.file.core.windows.net` plus `*.azure.com`); auth is identity-only — embedded keys do not work. Read-only except Azure Files.
- **Reserved mount paths**: `/dev`, `/dev/log`, `/tmp`, `/var`, `/var/log`. Max 15 volumes; no path may be a parent of another.
- `filesystemGroupId` defaults to 0 (root) when unset — set it (1-65534) for a non-root app.
Volume sets: `stateful-storage` skill.
### Service-to-service failures
**Symptoms:** a workload can't reach another internally (connection refused or timeout).
- **Target's internal firewall** must allow the caller — `same-gvc` / `same-org` / `workload-list` (default `none`); listing a workload needs `view` on it.
- **Endpoint** — `http://WORKLOAD.GVC.cpln.local:PORT` (use `http://`; the sidecar adds mTLS). An omitted port defaults to the target's first container port; only listed ports are reachable.
- **Cross-GVC** — the target must allow `same-org` or list the caller; traffic may span locations and incur egress charges.
### Dedicated load balancer & domain
**Symptoms:** unreachable after enabling a dedicated LB, TCP broken, wrong Host header.
- **TCP** needs a dedicated LB on the GVC plus a custom Domain with a TCP port — not on default endpoints (HTTP/HTTP2/gRPC only).
- Enabling/disabling a dedicated LB causes brief DNS-propagation downtime and per-location charges.
- **Serverless Host header** — a custom domain delivers the canonical endpoint as `Host` (the custom domain moves to `X-Forwarded-Host`); standard and stateful workloads get the custom domain as `Host`.
- **Protocol compatibility** — the domain port protocol must match the container's: HTTP2 fronts HTTP2 or gRPC, HTTP fronts HTTP.
Routing, TLS, and LBs: `domain` and `ipset-load-balancing` skills.
## Apply fixes within the schema's limits
A fix the Joi schema rejects at `update_workload` time is worse than none. Before applying a resource or option change, confirm it stays within these (the validator names the violated rule):
- `memory` (MiB) at most 8× `cpu` (millicores); CPU ≥ 25m; memory ≥ 32MiB; `minCpu`/`minMemory` never above `cpu`/`memory`.
- `runAsUser` / `filesystemGroupId`: 1-65534 (0/root rejected).
- `terminationGracePeriodSeconds`: ≤ 900 (higher only with the `cpln/relaxGracePeriodMax` tag).
- `capacityAI` with `metric: cpu` is rejected; `capacityAI` is stripped on stateful, cron, and vm.
- Standard/stateful `minScale: 0` requires `metric: keda`; cron and vm cannot scale to zero.
- A metric outside the workload type's allow-list is rejected.
Prefer `update_workload` (PATCH — only the fields you set change). For full manifest control, author against `get_resource_schema` then `cpln apply -f workload.yaml --gvc GVC`.
## Verify
After applying, poll `list_deployments` until every location reports ready, confirm the original symptom cleared (events/logs), and report the **canonical endpoint** `list_deployments` returns — never a constructed URL. For a public workload, confirm it actually responds, not just that it is ready.
## Troubleshooting
| Symptom | Likely cause | First check |
|---|---|---|
| Restarts; `OOMKilled` in events | memory cap too low (or Capacity AI downsized) | `query_metrics` memory; raise `memory` + `cpu` |
| `ImagePullBackOff` / stuck | bad image ref, wrong platform, missing pull secret | events; GVC `pullSecretLinks` |
| Env vars empty | broken identity + `reveal` chain or wrong reference | `grant_workload_secret_access` |
| Healthy but 502/503 | spec port ≠ listening port, or a blocked port | workload spec and startup logs |
| Unreachable / can't call out | deny-by-default firewall | `firewallConfig` |
| Won't become ready | probe path/port wrong or too aggressive | events; probe config |
| Can't reach another workload | target internal firewall `none`, or `https://` used | target `inboundAllowType`; use `http://` |
| 502/503 during deploys | missing `sleep`, or app ignores SIGTERM | `which sleep`; SIGTERM handling |
| `update_workload` rejected | the fix violates a schema limit | *Apply fixes within the schema's limits* |
## Quick reference
### MCP tools
| Tool | Purpose |
|---|---|
| `list_deployments` | Primary per-location readiness monitor |
| `get_workload_events` | Image / crash / probe / schedule events |
| `get_workload_logs` | App and `_accesslog` logs |
| `list_workload_replicas` | List running replicas |
| `query_metrics` (after `list_metrics`) | Memory / CPU / latency pressure |
| `update_workload` | Apply a spec fix (PATCH) |
| `grant_workload_secret_access` | Build the identity + `reveal` chain in one call |
| `get_resource_schema` | Exact fields before a manifest-level fix |
### CLI (fallback)
Use when MCP is unavailable, for an interactive shell, or in CI/CD (service-account `CPLN_TOKEN`).
```bash
cpln workload get WORKLOAD --gvc GVC -o json
cpln workload eventlog WORKLOAD --gvc GVC -o json
cpln workload connect WORKLOAD --gvc GVC --location LOCATION
cpln apply -f workload.yaml --gvc GVC
```
### Related skills
- `workload` — primary skill (types, spec shape, tool division); start here.
- `workload-security` — probe tuning, termination, `securityOptions`, direct LBs.
- `firewall-networking` — inbound / outbound / internal rules.
- `autoscaling-capacity` — scaling strategies and Capacity AI.
- `access-control` / `setup-secret` — the identity + `reveal` chain.
- `stateful-storage` — volume sets; `domain` / `ipset-load-balancing` — routing and LBs; `image` — registries and pull secrets.
## Documentation
- [Workload Types](https://docs.controlplane.com/reference/workload/types.md)
- [Containers](https://docs.controlplane.com/reference/workload/containers.md)
- [Capacity AI](https://docs.controlplane.com/reference/workload/capacity.md)
- [Firewall](https://docs.controlplane.com/reference/workload/firewall.md)
- [Termination](https://docs.controlplane.com/reference/workload/termination.md)
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Control Plane Corporation
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 12:00 UTC
- Collection status
- Collected
plugin_asdk_app_6a3345aed5b081918ae752ac49e4df0e
Download plugin data (JSON)