← Control PlaneCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Control Plane
Snapshot Sep 30, 2026 · 23:00 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "logql-observability",
"description": "Queries workload logs with LogQL on Control Plane. Use to troubleshoot a workload from its logs, or for log search, access or egress logs, cron run logs, missing or truncated logs, retention, or logs in Grafana.",
"included_files": [],
"skill_md_contents": "---\nname: logql-observability\ndescription: \"Queries workload logs with LogQL on Control Plane. Use to troubleshoot a workload from its logs, or for log search, access or egress logs, cron run logs, missing or truncated logs, retention, or logs in Grafana.\"\n---\n\n# LogQL & Log Observability\n\nControl Plane stores workload stdout/stderr in Loki and queries it with LogQL. The org is the Loki tenant — it comes from the endpoint path, so `org` is never a query label and queries cannot cross orgs. Reading logs requires the org-level `readLogs` permission, and the in-pod `CPLN_TOKEN` cannot authenticate to the logs endpoint — use a user or service-account token (see the `workload` skill). The recurring agent failure is passing a raw `query` to the MCP tool alongside structured params: a raw query replaces them entirely (the tool rejects the combination), so a raw query must embed every label itself.\n\n## Two ways to query\n\n- **MCP (primary for agents):** `get_workload_logs` — structured params `gvc` (required), `workload`, `container`, `location`, `filter` (literal substring, LogQL `|=`, not regex); window `since` (default `1h`) or `from`/`to` (ISO 8601, `from` inclusive, `to` exclusive); `limit` (default 30, max 999, single request — `truncated: true` means narrow the window or filter harder); `order` (`oldest_first` default, or `newest_first`). For regex, parsers, or the `replica`/`stream`/`version` labels, pass a raw `query` (max 500 chars) instead of the structured selectors.\n- **CLI:** `cpln logs '<LOGQL>'` — interactive debugging (live `--tail`), scripts, CI.\n\n```bash\n# Defaults: --since 1h, --limit 30, --direction forward\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"}' --org ORG\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"} |= \"error\"' --limit 100\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\", container=\"main\"}' --since 7d\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"}' --tail # live follow; the server ends a tail session after 30m\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"}' \\\n --from 2026-06-01T00:00:00Z --to 2026-06-02T00:00:00Z # ISO 8601 or relative (7d, now-1M); from inclusive, to exclusive\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"} |= \"error\"' --since 24h --limit 0 # 0 = unlimited, auto-paginates\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\"}' -o jsonl # one JSON object per line; -o raw = bare lines\n```\n\n- `--since` takes relative durations (`ms s m h d w mo y`, compound like `1h30m`). `--from`/`--to` (CLI) take an ISO 8601 timestamp **or** a relative duration meaning that long ago — `--from 7d`, `--from now-1M`, `--to now-30m` (the CLI accepts `M` for months; the `get_workload_logs` tool takes ISO 8601 only).\n- `cpln workload eventlog WORKLOAD` (alias `cpln workload log`) is resource event history, not container output — for application logs always use `cpln logs`.\n\n## Labels\n\n| Label | Value |\n|:---|:---|\n| `gvc` | GVC name |\n| `workload` | Workload name |\n| `container` | Container name, or a built-in stream below |\n| `location` | Deployment location, e.g. `aws-us-east-1` |\n| `provider` | Cloud provider |\n| `replica` | Replica (pod) name — unique per cron execution |\n| `stream` | `stdout` or `stderr` |\n| `version` | Workload deployment version that wrote the line |\n\nAt least one non-empty matcher is required; regex matchers work — `{gvc=~\".+\"}` spans every GVC in the org.\n\n## Filters and LogQL features\n\n| Operator | Meaning | Example |\n|:---|:---|:---|\n| `\\|= \"text\"` | contains | `\\|= \"error\"` |\n| `!= \"text\"` | does not contain | `!= \"health\"` |\n| `\\|~ \"regex\"` | matches regex | `\\|~ \"timeout\\|crash\"` |\n| `!~ \"regex\"` | does not match | `!~ \"debug\\|trace\"` |\n\nLoki is current (3.x), so full LogQL works: parsers (`| json`, `| logfmt`, `| pattern`, `| regexp`), post-parse label filters (`| latency > 100`), `line_format`, and metric queries (`count_over_time`, `rate`, `sum ... by`). The CLI and MCP tool print log lines only — run metric queries in Grafana: the `Explore on Grafana` link on the console Logs page opens it with the query prefilled (the org `grafanaAdmin` permission grants the Grafana Admin role; everyone else is Viewer).\n\n```logql\n{gvc=\"GVC\", workload=\"WORKLOAD\"} |= \"error\" != \"health\" # errors minus noise\n{gvc=\"GVC\", workload=\"WORKLOAD\"} |~ \"panic|fatal|exception\" # crashes and stack traces\n{gvc=\"GVC\", workload=\"WORKLOAD\", container=\"_accesslog\"} |= \"\\\" 50\" # HTTP 5xx in access logs\nsum(count_over_time({gvc=\"GVC\", container=\"_accesslog\"} |= \"\\\" 50\"[1m])) by (workload) # 5xx rate (Grafana)\n```\n\n## Built-in log streams\n\n| Selector | Contents |\n|:---|:---|\n| `container=\"_accesslog\"` | Inbound requests (Envoy access-log format) on the workload's ports |\n| `container=\"_requestlog\"` | The workload's outbound (egress) requests through the sidecar, same format |\n| `workload=\"_loadbalancer\"` | Access logs of the GVC's dedicated load balancer |\n| `container=\"_alerts\"` | Threat-detection (Falco) alerts, with extra labels `rule`, `priority`, `source` |\n\nPlatform health probes and unroutable-request noise are filtered out of `_accesslog` by design, so probe traffic never shows up there. Sidecar and system containers (`istio-init`, `istio-validation`, `cpln-*`, `debugger-*`) are never collected.\n\n## Why logs go missing (pipeline limits)\n\n- **Lines over 16 KiB are cut** at 16 KiB with a `... [truncated N bytes]` suffix; empty lines are dropped.\n- **Per-replica rate limit:** each replica+container pair is capped (roughly 10,000 lines/s on managed clusters; effectively unlimited on BYOK). Excess lines are dropped, and when collection resumes one marker line appears: `#### Replica logs were rate-limited: N lines in the last Xs were not collected ####`.\n- **Retention:** org spec `observability.logsRetentionDays`, default 30 (0 turns log collection off entirely; the `org-management` skill owns org spec edits). Queries beyond retention return nothing, not an error.\n- **Server caps:** one query may span at most 31 days and times out after 2 minutes; tail sessions end after 30 minutes.\n\n## Cron workloads: logs for one execution\n\n`{gvc=, workload=}` on a cron workload interleaves every past run. Each execution runs in its own replica, so scope to one run with the `replica` label plus the execution's time window.\n\n**1. List executions.** `list_deployments` **with** `location` returns the full deployment JSON; `status.jobExecutions[]` holds the per-execution metadata (without `location` the tool returns only a readiness summary). CLI:\n\n```bash\ncpln workload get-deployments WORKLOAD --gvc GVC -o json | jq '\n [(.items // .)[] as $d | $d.status.jobExecutions[]? | . + {location: $d.name}]\n | sort_by(.startTime)[] | {location, name, status, startTime, completionTime, replica}'\n```\n\nSchema facts (nodelibs `cronjob.ts`) that bite scripts:\n\n- `status` is one of `successful | failed | active | pending | invalid | removed` (default `pending`). Running means `status == \"active\"` AND no `completionTime`; a missing `completionTime` alone proves nothing (the schema warns it is not an indication of success).\n- `completionTime: null` is stripped server-side and `startTime` may be absent for never-started runs — each field is present-with-value or absent, so guard flags with `${VAR:+--from \"$VAR\"}` and jq `// empty`.\n- `replica` is optional: when missing, the run never got a pod and there are **no logs** — diagnose with the execution's `containers` map and `message` field (aggregated pod events) and `get_workload_events`.\n\n**2. Query that replica, time-bounded.** MCP: raw `query` embedding ALL labels (structured params must be omitted) plus `from`/`to`. CLI:\n\n```bash\ncpln logs '{gvc=\"GVC\", workload=\"WORKLOAD\", location=\"LOCATION\", replica=\"REPLICA-ID\"}' \\\n --org ORG --from START_TIME --to COMPLETION_TIME\n```\n\n`jobExecutions` timestamps are ISO 8601 and pass straight through. Pad the window a minute or two on each side, and widen it before concluding logs do not exist. For a live run (`active`, no `completionTime`), drop `--to` and add `--tail`.\n\n## Quick reference\n\n| Tool | Use |\n|:---|:---|\n| `get_workload_logs` | LogQL queries — structured params or raw `query` |\n| `list_deployments` | Deployment health; cron `status.jobExecutions` (pass `location`) |\n| `get_workload_events` | Probe failures, scheduling, restarts — events, not app logs |\n\nCI/CD and headless use: set `CPLN_TOKEN` and run `cpln logs` directly (no profile needed); the principal must hold org `readLogs`.\n\n## Troubleshooting\n\n| Symptom | Cause and fix |\n|:---|:---|\n| 403 `requires permission \"readLogs\" in org` | Grant `readLogs` on the org via policy (it implies `view`) |\n| `Invalid --from format: ...` | `--from`/`--to` accept an ISO 8601 timestamp or a relative duration (`7d`, `now-1M`); correct the value |\n| Empty result for known activity | Window outside retention (default 30d), span over 31 days, or filters too narrow; app may not write to stdout/stderr |\n| `#### Replica logs were rate-limited ... ####` marker | Per-replica cap was hit; lines in that interval are gone — reduce log volume |\n| Line ends with `... [truncated N bytes]` | 16 KiB per-line cap — emit smaller lines |\n| Health checks absent from `_accesslog` | Filtered by design; probe failures surface in `get_workload_events` |\n| MCP rejects raw `query` combined with `workload` etc. | A raw query replaces the structured params — embed all labels in the query itself |\n| `cpln workload log` shows no app output | That is the eventlog alias; use `cpln logs` |\n\n## Related skills\n\n| Skill | Owns |\n|:---|:---|\n| `workload` | Deploy and diagnose flow, injected `CPLN_*` env vars, canonical URLs |\n| `metrics-observability` | PromQL, default metrics, Grafana alert rules, Prometheus federation |\n| `external-logging` | Shipping logs to S3, Datadog, Coralogix, and other providers |\n| `mk8s-byok` | mk8s cluster logs add-on (`cluster_name` / `namespace` labels) |\n\n## Documentation\n\n- [Logs Reference](https://docs.controlplane.com/core/logs.md)\n- [CLI logs Command](https://docs.controlplane.com/cli-reference/commands/logs.md)\n- [External Logging Overview](https://docs.controlplane.com/external-logging/overview.md)\n- [LogQL (upstream Grafana reference)](https://grafana.com/docs/loki/latest/query/)\n"
}SHA-256: a6ddd220c72d78f8cdb23c6310cbf59a0f6baef0861ba32f8e6b5d80ce7daa36