{"id":12100,"plugin_id":"plugin_asdk_app_6a3345aed5b081918ae752ac49e4df0e","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:00:54.965Z","digest":"87321f1ed80f7f385c0a17bbc7edaf99b16761f3245afd2d4358755027c2e70f","against":null,"payload":{"name":"metrics-observability","description":"Workload metrics, PromQL, Grafana, and tracing on Control Plane. Use to observe or troubleshoot a workload via CPU/memory/request or custom metrics, traces, Prometheus federation, alerts, or metrics retention.","included_files":[],"skill_md_contents":"---\nname: metrics-observability\ndescription: \"Workload metrics, PromQL, Grafana, and tracing on Control Plane. Use to observe or troubleshoot a workload via CPU/memory/request or custom metrics, traces, Prometheus federation, alerts, or metrics retention.\"\n---\n\n# Metrics, Tracing & Observability\n\nControl Plane stores every workload's metrics as Prometheus-compatible time series in a managed backend (Mimir), queryable in PromQL through the per-org managed Grafana or the MCP tools. The org is the tenant — it comes from the endpoint path, so there is no `org=` label and no cross-org queries. Two traps dominate. **Series names are short:** memory is `mem_used` / `mem_reserved` / `mem_billable`, not `memory_*` — a `memory_used` query returns nothing, so ground names with `list_metrics` first. **Rate-shaped metrics are pre-rated:** `egress`, `requests_per_second`, and the latency buckets are already rated by the platform's recording rules, so you query them bare — wrapping them in `rate()` again returns garbage. Finally, a workload's in-pod `CPLN_TOKEN` cannot authenticate to the metrics endpoint; querying from outside the mesh needs a user or service-account token.\n\n## Two ways to query\n\n- **MCP (primary for agents):** `query_metrics` runs a PromQL query — a range query over the last `1h` at `60s` step by default; pass `resolution: \"instant\"` for a single point, or `since` / `from` / `to` / `step` to adjust. `list_metrics` discovers the metric names and real label values present in the org right now (built-in, `kube_`/`node_`, and custom); pass `metric:` to ground one metric's live labels before filtering. Reach for it whenever a query returns no series. Measure first, then change scaling settings.\n- **Grafana:** the managed per-org instance — open **Metrics** in the Console sidebar (or the **Metrics** link on any workload), use **Explore** for ad-hoc PromQL, and dashboards/alerting for the rest. The `grafanaAdmin` org permission grants the Grafana Admin role; everyone else is Viewer.\n\n`list_metrics`' built-in catalog still spells memory `memory_*`; trust the live names it returns (and this skill) — the queryable series is `mem_*`.\n\n## PromQL: query the right shape\n\nThe platform pre-computes rates, so the shape decides the query form:\n\n- **Gauges — query bare:** `cpu_used`, `mem_used`, `replica_count`, `workload_ready_replicas`.\n- **Pre-rated gauges — query bare, never `rate()`:** `egress` and `cross_zone_traffic` (bytes per minute), `requests_per_second`, `requests_initiated_per_second`, `cron_execution_rate`.\n- **Histogram — `histogram_quantile`, no extra `rate()`:** `request_duration_ms_bucket` keeps its `le` label and is already rated.\n- **Cumulative counters — wrap in `increase()` / `rate()` for velocity:** `container_restarts`, `cron_executions`, `workload_progress_failure`, `workload_rescheduled_replicas`, `domain_warnings`.\n\n```promql\ncpu_used                                              # cores in use, per replica (bare gauge)\nsum by (workload) (mem_used)                          # memory bytes per workload — mem_, not memory_\negress                                                # outbound bytes/minute (already rated — no rate())\nsum by (workload) (requests_per_second{response_class=\"500\"})   # 5xx rate; response_class is \"200\"..\"500\"\nhistogram_quantile(0.95, sum by (le) (request_duration_ms_bucket))   # p95 latency (ms); no rate() wrapper\nsum by (gvc, workload) (increase(container_restarts[5m]))           # restarts in the last 5m\n```\n\n## Built-in metrics\n\nCollected for every workload, no configuration. Names and types below are the recording-rule outputs (the queryable series). Call `list_metrics` for the complete live set, including your custom metrics.\n\n**Resource & network** (per replica): `cpu_used` / `cpu_reserved` / `cpu_billable` (cores, gauge); `mem_used` / `mem_reserved` / `mem_billable` (bytes, gauge); `egress` / `cross_zone_traffic` (bytes/minute, pre-rated gauge); `replica_count` / `workload_ready_replicas` / `workload_desired_replicas` (gauge).\n\n**Traffic** (per pod): `requests_per_second` and `requests_initiated_per_second` (pre-rated gauge, label `response_class`); `request_duration_ms_bucket` (latency histogram, keeps `le`).\n\n**Stability**: `container_restarts`, `workload_progress_failure`, `workload_rescheduled_replicas`, `cron_executions`, `domain_warnings` (cumulative counters); `cron_execution_rate` (pre-rated); `capacity_ai_updates`, `load_balancer` (gauge).\n\n**Volume** (per volume set): `volume_set_capacity_bytes`, `volume_set_used_bytes`, `volume_set_free_bytes`, `volume_set_billable_bytes`, `volume_set_capacity_billable`, `volume_set_snapshots_billable`.\n\n**Org-wide** (no workload label): `logs_storage_mb` / `metrics_storage_mb` / `tracing_storage_mb`; `agent_peers_count` / `agent_services_count` (gauge) and `agent_{tx,rx}_{bytes,packets}_total` (counter) from wormhole agents; `threat_detection_alerts` / `threat_detection_forward_total` / `threat_detection_forward_enabled`.\n\nmk8s clusters with metrics enabled also expose `kube_*` (kube-state-metrics) and `node_*` (node-exporter).\n\n## Custom metrics\n\nA container exposes Prometheus-format metrics by declaring a `metrics` block; the platform scrapes every replica every **30 seconds** (5s timeout). Set it at creation with `create_workload` or add it later with `update_workload`; if the typed tool doesn't surface the nested field, fall back to `get_resource_schema` for `workload` then `cpln apply -f workload.yaml`.\n\n```yaml\nspec:\n  containers:\n    - name: app\n      metrics:\n        port: 9100          # required; ≥80 and NOT a reserved port (see trap below)\n        path: /metrics      # required; string, max 128, default /metrics\n        dropMetrics:        # optional; RE2 regexes, dropped before scrape\n          - '^go_.*'\n          - '^process_.*'\n```\n\n- **Reserved-port trap:** `port` rejects the platform's sidecar ports — `9090`, `9091`, `8012`, `8022`, `15000`/`15001`/`15006`/`15020`/`15021`/`15090`, `41000`. The obvious Prometheus default `9090` fails; use `9100`, `2112`, etc.\n- Metric names starting with `cpln_` are dropped (you cannot overwrite platform series).\n- Scraped samples gain labels `org`, `gvc`, `workload`, `container`, `location`, `provider`, `region`, `cluster_id`, `replica`.\n\n## Distributed tracing\n\nTracing answers a different question than metrics: not \"is latency high?\" but **where** in the request path. It is opt-in via `spec.tracing` on a **GVC** (or org-wide on the org spec) — set it with `update_gvc` / `create_gvc` or `cpln apply`. Exactly one provider (`.xor`), and `sampling` (a required `0`–`100` percentage):\n\n- **`controlplane`** — built-in backend, queryable with the tools below; zero extra infrastructure.\n- **`otel`** — ship spans to your own OpenTelemetry collector (`endpoint`).\n- **`lightstep`** — ship to Lightstep (`endpoint` + an opaque `credentials` secret).\n\n`customTags` adds fixed key/values to every span (each value max 50 chars). Only requests served after enablement, in the sampled fraction, produce traces. Apps wanting to emit their own spans to the `controlplane` provider send OTLP to `tracing.controlplane:80` (gRPC) or `tracing.controlplane:4318` (HTTP).\n\nQuery the built-in backend with `query_traces` — structured params (`gvc`, `workload`, `location`, `errorsOnly`, `minDuration: \"500ms\"`) or a raw `traceql` query that replaces them; span attributes are `resource.gvc` / `resource.workload` / `resource.location`. Then `get_trace` reads one trace's span tree to name the slow/failed span. **Empty results are usually configuration:** confirm tracing is enabled, sampling catches traffic, and the window saw requests. Triage flow: `query_traces` (`minDuration` or `errorsOnly`) to the worst trace, `get_trace` to the culprit span, then `get_workload_logs` over the same window for the application error.\n\n## Built-in Grafana alert rules\n\nThe managed Grafana provisions five rules, all annotated to the `cpln-metrics-overview` dashboard. They evaluate on import but deliver nothing until a Grafana contact point exists — set `defaultAlertEmails` (below) or add a contact point. Deletions are recreated on next login.\n\n| Rule | Fires when | Default |\n|:---|:---|:---|\n| `container-restarts` | `increase(container_restarts[5m]) > 0` per gvc/location/workload (any restart) | active |\n| `stuck-deployments` | more than one deploy `version` of a workload is restarting (15m) | active |\n| `workload-progress-failure` | `increase(workload_progress_failure[10m]) > 0` (15m) | active |\n| `threat-detection-alerts` | `increase(threat_detection_alerts[15m]) > 0` per gvc/workload/priority/rule | active |\n| `domain-warnings` | `increase(domain_warnings[60m]) > 5` per domain/type | **paused** |\n\n## Retention & billing\n\nRetention and default alert recipients live in the org `observability` block. No typed MCP tool edits it (the `org-management` skill owns org-spec edits) — apply via CLI: `get_resource_schema` for `org`, then `cpln apply -f org.yaml`.\n\n```yaml\nkind: org\nspec:\n  observability:\n    logsRetentionDays: 30       # int 0-3650, default 30 (0 disables log collection)\n    metricsRetentionDays: 30    # int 0-3650, default 30\n    tracesRetentionDays: 30     # int 0-3650, default 30\n    defaultAlertEmails:         # email[]; recipients for the grafana-default-email contact point\n      - ops@example.com\n```\n\nCombined storage of logs, metrics, and traces is charged per GB-month over 100 GB.\n\n## Export & centralize metrics\n\n`readMetrics` (org permission, \"access usage and performance metrics\") gates both the federation endpoint and Grafana data sources. Create a service account and grant it `readMetrics` via policy — the `access-control` skill owns that; here is the metrics-specific wiring.\n\n**Federate into your own Prometheus** — scrape the source org with the service-account token:\n\n```yaml\nscrape_configs:\n  - job_name: cpln-federate\n    scheme: https\n    honor_labels: true\n    metrics_path: '/metrics/org/SOURCE_ORG/api/v1/federate'\n    params:\n      'match[]': ['{__name__=~\".+\"}']     # narrow the matcher to limit egress\n    authorization: { type: Bearer, credentials: \"${CPLN_SERVICE_ACCOUNT_TOKEN}\" }\n    static_configs:\n      - targets: ['metrics.cpln.io']\n```\n\n**Cross-org Grafana** — in a viewer org's Grafana, add a Prometheus data source with URL `https://metrics.cpln.io/metrics/org/SOURCE_ORG` and a custom HTTP header `authorization` = `Bearer <SOURCE_ORG_SA_TOKEN>`, then **Save & Test**. The community dashboard `grafana.com/dashboards/20378` (Multi-Source Metrics Overview) visualizes several at once.\n\n**Token trap:** `metrics.cpln.io` authenticates user and service-account tokens only. A workload's injected `CPLN_TOKEN` does **not** work there even with `readMetrics` on its identity — the metrics proxy forwards only the link headers, never the signed header the in-mesh API path injects (see the `workload` skill). Query from inside a workload with a service-account key.\n\n## Autoscaling metric availability\n\nThis skill covers only which scaling metrics each workload type allows; for strategy, YAML, percentiles, multi-metric, KEDA, and Capacity AI, see the `autoscaling-capacity` skill.\n\n| Metric | Serverless | Standard | Stateful |\n|:---|:---:|:---:|:---:|\n| `concurrency` | yes | no | no |\n| `cpu` / `memory` / `rps` | yes | yes | yes |\n| `latency` / `keda` | no | yes | yes |\n| `disabled` | yes | yes | yes |\n\n`vm` workloads allow only `disabled`; `cron` has no autoscaling. (`memory` here is the scaling keyword — distinct from the `mem_used` series.)\n\n## Quick reference\n\n| Tool | Use |\n|:---|:---|\n| `list_metrics` | Discover real metric names and label values (built-in + custom) before querying |\n| `query_metrics` | Run a PromQL query against the org's metrics |\n| `query_traces` | Search traces (TraceQL) — slow (`minDuration`) or failed (`errorsOnly`) requests |\n| `get_trace` | Read one trace's span tree to locate the slow/failed span |\n| `get_workload_logs` | Correlate a metric spike with logs (see `logql-observability`) |\n\n- **Metrics endpoint:** `https://metrics.cpln.io/metrics/org/{ORG}` (federation adds `/api/v1/federate`).\n- **Permission:** `readMetrics` (federation endpoint + Grafana data source).\n- No typed tool edits the org `observability` block, the GVC `tracing` block, or a container `metrics` block — fall back to `get_resource_schema` + `cpln apply`.\n\n## Troubleshooting\n\n| Symptom | Cause and fix |\n|:---|:---|\n| Query returns no series | Wrong name — memory is `mem_used`, not `memory_used`; run `list_metrics` to confirm live names |\n| `egress`/latency values look tiny or wrong | Pre-rated series wrapped in `rate()` — query `egress` bare, latency via `histogram_quantile(..., request_duration_ms_bucket)` |\n| Custom `metrics` block rejected | `port` is reserved (`9090`/`9091`/`15000`+) or below 80 — use `9100`/`2112` |\n| Custom metrics never appear | Names prefixed `cpln_` are dropped; scrape runs every 30s — allow a cycle |\n| 403 at `metrics.cpln.io` | Principal lacks `readMetrics`, or an in-pod `CPLN_TOKEN` was used — use a user/SA token |\n| `query_traces` empty | Tracing not enabled on the GVC, sampling too low, or no traffic in the window |\n| Alert never notifies | Rules evaluate but need a contact point — set `defaultAlertEmails` or add one in Grafana (`domain-warnings` also ships paused) |\n\n## Related skills\n\n| Skill | Owns |\n|:---|:---|\n| `workload` | Deploy/diagnose flow, injected `CPLN_*` env vars, the spec that holds `metrics` |\n| `autoscaling-capacity` | Scaling strategy, per-metric YAML, percentiles, KEDA, Capacity AI |\n| `logql-observability` | Log queries (LogQL), `cpln logs`, correlating spikes with log events |\n| `org-management` | Org-spec edits — the `observability` retention block |\n| `external-logging` | Shipping logs to S3, Datadog, Coralogix, and other providers |\n\n## Documentation\n\n- [Default Metrics](https://docs.controlplane.com/guides/default-metrics.md)\n- [Custom Metrics](https://docs.controlplane.com/reference/workload/custom-metrics.md)\n- [Export Metrics (federation)](https://docs.controlplane.com/guides/export-metrics.md)\n- [Centralized Metrics](https://docs.controlplane.com/guides/centralized-metrics-management.md)\n- [Autoscaling](https://docs.controlplane.com/reference/workload/autoscaling.md)\n- [PromQL (upstream Prometheus reference)](https://prometheus.io/docs/prometheus/latest/querying/basics/)\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}