{"id":19430,"plugin_id":"plugins_6a992fa57b6481918dffd35d23f03408","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:15:34.447Z","digest":"673f9878ffbb08009ab4f48426607229e31cc01dbd366c1d5100a2e04ce5519c","against":null,"payload":{"description":"Use when the user reports that a Zilliz Cloud cluster or Milvus collection is unhealthy, slow, stuck, returning errors, hitting quotas, or otherwise misbehaving — or when they ask \"what's wrong with...\", \"why is ... slow\", \"diagnose ...\", \"troubleshoot ...\".","included_files":[],"name":"diagnose","skill_md_contents":"---\nname: diagnose\ndescription: Use when the user reports that a Zilliz Cloud cluster or Milvus collection is unhealthy, slow, stuck, returning errors, hitting quotas, or otherwise misbehaving — or when they ask \"what's wrong with...\", \"why is ... slow\", \"diagnose ...\", \"troubleshoot ...\".\n---\n\n## Prerequisites\n\n1. CLI installed and logged in (see setup skill).\n2. For cluster diagnosis: a cluster context, or pass `--cluster-id` explicitly.\n3. For collection diagnosis: the cluster context must point at the cluster that owns the collection (see setup skill).\n\n## Scope\n\nThis skill performs **read-only diagnosis** using existing `zilliz` commands. It does NOT mutate clusters, collections, indexes, or data. Every recommendation is presented as a suggestion plus the exact next command the user can run themselves.\n\nTwo entry points:\n\n| User says | Use section |\n|---|---|\n| \"diagnose my cluster\", \"cluster is slow / stuck / unhealthy\" | [Cluster Diagnosis](#cluster-diagnosis) |\n| \"diagnose this collection\", \"search is slow on X\", \"X won't load\" | [Collection Diagnosis](#collection-diagnosis) |\n\nWhen the user's intent spans both (e.g. \"everything is slow\"), run cluster diagnosis first — a collection-level symptom often has a cluster-level cause.\n\n## Cluster Diagnosis\n\nFollow this checklist in order. Stop early only if you find a P0 problem (cluster not RUNNING, quota hard-stop) — surface it before running the rest.\n\n### 1. Collect\n\nRun these in parallel where possible and prefer `-o json` for machine parsing:\n\n```bash\nzilliz context current\nzilliz cluster describe --cluster-id <id> -o json\nzilliz cluster list --all -o json                       # peer clusters for context\nzilliz billing usage -o json                            # quota / spend headroom\n```\n\nThen pull time-series metrics covering the last hour and last 24h. Use the chart output for human review and `-o json` if you need to compute thresholds:\n\n```bash\nzilliz cluster metrics --cluster-id <id> \\\n  -m CU_COMPUTATION -m CU_CAPACITY -m CU_SIZE -m REPLICA_COUNT \\\n  -m STORAGE -m SLOW_QUERIES \\\n  -m SEARCH_QPS -m SEARCH_LATENCY_P99 \\\n  -m SEARCH_FAIL_RATE -m INSERT_FAIL_RATE \\\n  --period 1h\n```\n\nRepeat with `--period 24h` for trend context.\n\n### 2. Analyze\n\nWalk each rule. Each finding must cite the evidence (command + observed value).\n\n| Check | Evidence | If true, suggest |\n|---|---|---|\n| Cluster status ≠ RUNNING | `cluster describe` `.status` | Pause diagnosis; explain status; if SUSPENDED suggest `cluster resume` |\n| Plan vs. observed QPS mismatch (e.g. Serverless under sustained high QPS) | plan + `SEARCH_QPS` | Discuss plan upgrade tradeoff |\n| `CU_COMPUTATION` p95 > 80% of `CU_CAPACITY` for sustained windows | metric series | Recommend scaling CU or adding replicas |\n| `SEARCH_LATENCY_P99` rising while QPS flat | latency vs qps | Likely index/CU pressure; drill into top collections |\n| Non-zero `*_FAIL_RATE` | fail-rate series | Cross-check with recent user-reported errors |\n| `SLOW_QUERIES` non-zero | metric | Drill into per-collection diagnosis for the offenders |\n| Storage trending toward quota | `STORAGE` + billing | Suggest cleanup / plan change before hard cap |\n| Region likely far from client | endpoint + user-reported RTT | Note possible region mismatch — cannot measure RTT from CLI |\n\n### 3. Present\n\nRender a single report with three sections, in this order:\n\n1. **Summary** — one-line health verdict (healthy / degraded / critical) + 1-3 sentence rationale.\n2. **Findings** — table with columns `Severity | Finding | Evidence | Suggested Next Step`.\n3. **Suggested commands** — copy-pasteable `zilliz ...` commands matching each suggestion. Never run mutating commands yourself — let the user run them.\n\n## Collection Diagnosis\n\n### 1. Collect\n\n```bash\nzilliz collection describe --name <coll> -o json\nzilliz collection get-stats --name <coll> -o json\nzilliz collection get-load-state --name <coll> -o json\nzilliz index list --collection-name <coll> -o json\n# For each vector index reported above:\nzilliz index describe --collection-name <coll> --field-name <field> -o json\nzilliz partition list --collection-name <coll> -o json\n```\n\nPull per-collection metrics for the last hour and 24h:\n\n```bash\nzilliz collection metrics -c <coll> \\\n  -m SEARCH_QPS -m SEARCH_LATENCY_P99 -m SEARCH_FAIL_RATE \\\n  -m QUERY_QPS -m QUERY_LATENCY_P99 \\\n  -m INSERT_QPS -m INSERT_LATENCY_P99 \\\n  -m ENTITIES -m ENTITIES_LOADED -m ENTITIES_INDEXED \\\n  --period 1h\n```\n\nIf the cluster context is wrong or the collection lives in a non-default database, pass `--database <db>` on every command.\n\n### 2. Analyze\n\n| Check | Evidence | If true, suggest |\n|---|---|---|\n| Load state ≠ Loaded (or partially loaded) | `get-load-state` | `collection load --name <coll>`; explain that unloaded collections cannot serve search |\n| `ENTITIES_LOADED` ≪ `ENTITIES` | metrics | Load incomplete or recently grew; wait or reload |\n| `ENTITIES_INDEXED` ≪ `ENTITIES` | metrics | Indexing lag; investigate before tuning |\n| Vector field present but no vector index, or FLAT on large row count | `index list` + `get-stats` | Recommend HNSW / IVF_* with a parameter range appropriate for row count; mark as recommendation, not absolute |\n| Index params clearly off (e.g. `nlist` ≪ √N for IVF) | `index describe` + row count | Suggest revised values; note that benchmarking is needed for the final number |\n| Replica count = 1 with sustained high QPS | `collection describe` replicas + `SEARCH_QPS` | Suggest more replicas, conditioned on CU headroom from cluster diagnosis |\n| Schema reasons: very wide varchar, many un-indexed scalar fields used in filters | `collection describe` schema | Note impact; suggest scalar indexes where applicable |\n| No partition key but row count is large and queries are naturally partitionable | schema | Mention partition-key option (cannot be added in place — design-time decision) |\n| Non-zero `SEARCH_FAIL_RATE` | metric | Cross-check with cluster-level fail-rate metrics |\n| Latency rising while QPS flat | latency vs qps | Index pressure or growing segment count; consider compaction; note that segment-level state is not exposed via CLI |\n\n### 3. Present\n\nSame three-section format as cluster diagnosis. When a collection finding's true root cause is at the cluster level (e.g. CU saturation), say so explicitly and reference the cluster report.\n\n## Guidance\n\n- **Read-only.** Never run `delete`, `drop`, `release`, `resume`, `suspend`, `create`, `update`, index create/drop, or any mutating data-plane command as part of diagnosis. Present them as suggestions only.\n- **Cite evidence.** Every finding must reference the command that produced it and the observed value. No unsourced claims.\n- **Mark uncertainty.** Index parameter recommendations, replica counts, and CU sizing depend on workload specifics the CLI cannot observe. Phrase as \"starting point, benchmark to confirm.\"\n- **Know the limits.** The CLI cannot see per-query plans, segment-level state, compaction backlog, GC, server-internal queues, or alert-rule state. If a question requires those, say so plainly rather than guessing.\n- **Prefer parallel collection.** When multiple read-only commands are independent, run them in parallel to keep the diagnosis fast.\n- **JSON for parsing, charts for humans.** Use `-o json` (optionally with `--query`) when comparing values against thresholds; use the default chart output when surfacing trends back to the user.\n- **Honor context.** If the user has multiple clusters, confirm which one before collecting. If the collection is in a non-default database, thread `--database` through every command.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}