← incident.ioCONTENT HISTORY

Update to incident.io

Snapshot Oct 9, 2026 · 12:28 UTC · version 1.20261009.616

Collection source: downloaded plugin package.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Query an organization's observability data — logs, metrics, traces, profiles, dashboards, Kubernetes, SQL databases — with the incident.io telemetry tools. Use when answering a question that needs evidence from its monitoring: what errors fired, when latency moved, which pod restarted, what a dashboard showed, what its databases hold. Not for questions the incident record already answers, and not for querying incident.io's own data.\n",
  "included_files": [],
  "name": "telemetry",
  "skill_md_contents": "---\nname: telemetry\ndescription: >\n  Query an organization's observability data — logs, metrics, traces, profiles, dashboards,\n  Kubernetes, SQL databases — with the incident.io telemetry tools. Use when answering a\n  question that needs evidence from its monitoring: what errors fired, when latency moved,\n  which pod restarted, what a dashboard showed, what its databases hold. Not for questions\n  the incident record already answers, and not for querying incident.io's own data.\n---\n\n# Telemetry\n\nAn organization's observability data lives in datasources it connected — a Loki, a\nPrometheus, a Honeycomb, a Postgres. Each holds different signals and speaks a different\nquery language. You reach all of them through two sets of tools: the native telemetry tools,\nand the extension connector tools where there are any. The next section says how to call\nthem; the rest of this skill applies however you call them.\n\n## How you reach telemetry\n\n### Native telemetry\n\n```\ntelemetry_guidance_show(path: \"data-sources.yaml\")\nlog_query(datasource_id: \"…\", purpose: \"…\", expression: \"…\", time_from: \"…\", time_to: \"…\")\n```\n\nThe telemetry tools are on the incident.io connection. Each tool's description is the\nauthority on it: what it does, the arguments it takes, and how to read what it returns.\n\nWhere the session has no `telemetry_guidance_show` but has `ask_telemetry`, the telemetry tools\nare not enabled for this organization: ask `ask_telemetry` the question in plain words\ninstead, and say so. Where it has neither, say the incident.io connection has no telemetry\ntools rather than guessing an answer.\n\n#### Running queries together\n\nRun queries that do not depend on each other as parallel tool calls, in one turn.\n\n### Extension connector\n\nConnector tools are on the incident.io connection, named `<connector>__<tool>`. If there are\nno such tools, there are no connectors.\n\n### Docs\n\nRead the datasource index, `data-sources.yaml`, and each file its `docs` fields list with\n`telemetry_guidance_show`, passing each path exactly as listed. Read each one whole, in one\nturn of parallel calls. If your client cuts a result short or saves it to a file, read on\nfrom where it stopped before you query.\n\n### Times\n\nPut the window in `time_from` and `time_to` as RFC3339: leave them out and the query covers\nonly the last hour. Convert anything a person expressed as a wall-clock time from the current\ndate in the conversation, rather than from your own sense of now.\n\n### Errors\n\nRead the error message, not the error code: the message says what went wrong. An internal\nerror with no detail means the call failed, not that the data is empty. A product error\nmeans the organization cannot query telemetry here — say so. A datasource that cannot be\nreached is a fact about the monitoring, not an answer to the question.\n\n## Connectors\n\nSome sources of telemetry data are made available through a connector rather than a native\ntelemetry data source. Connectors are not listed in `data-sources.yaml`.\n\nWhen picking where a signal lives, consider both the telemetry datasources and the\nconnectors, then query the one that holds it. Native telemetry wins when both hold the same\ndata; do not query both to confirm. Note that some vendors will be available through both\nnative telemetry (e.g. for logs and metrics) and a connector (e.g. for inspecting ingestion\nconfiguration). When the native datasource for a signal errors or returns nothing, check the\nconnectors before reporting the signal as unavailable.\n\nUse your judgement to decide whether a connector is appropriate to use to answer a telemetry\nquery - not all connectors, or tools within a connector, expose telemetry data. E.g. a\nconnector called `production operations` with a tool `read_service_restart_logs` would be\nappropriate for a telemetry query, whereas a connector called `linear` with a tool\n`fetch_project_updates` would not be.\n\n## The loop\n\n1. **Read `data-sources.yaml`**, then **the connectors**, to pick a datasource.\n   `data-sources.yaml` lists each one with its ID, type, what it holds and the earliest data\n   it still keeps — and, under `docs`, the paths of its query-language reference and its own\n   guidance. Match the question to a datasource that holds that signal: a metrics source\n   cannot answer a question about log text. Then check the window you need against that\n   datasource's earliest timestamp, rather than learning it from a refusal later.\n2. **Check `memory_recall` first**, where it is available. It returns expressions already\n   proven on this account for that datasource — adapt one, or replay it exactly with\n   `from_query_id`, rather than writing from scratch.\n3. **Read the tool's description**, and the connector tool's if you have decided to use a\n   connector, then run it.\n4. **Drill into what came back.** A telemetry query returns a `result_id`, and the drill-down\n   tools narrow that stored result without re-querying the source.\n5. **Cite what you used.** A finding without the query behind it cannot be checked.\n\n## Ask more than one question at a time\n\nQueries that do not depend on each other should run together rather than one after another,\nas [Running queries together](#running-queries-together) shows. An incident rarely turns on a\nsingle query, and running them in sequence spends an engineer's time for no reason.\n\nBatch different questions, not copies of one query you have not proven yet. A batch returns\nwhen its slowest query finishes, so it costs whatever its worst member costs, and a shape\nthat is too broad — or that a source has already turned down once — fails once per copy:\nslicing one unproven query across a dozen windows is a dozen ways to be told the same no.\nRun it once, read what came back, and fan out from there.\n\nA datasource serves every query from one shared budget. Queries that each read a lot — a\nselector the guidance says covers most of the data, a range of hours — slow each other down\nwhen they run together, and a batch of them can all time out where each alone would have\nfinished. Your first query against a selector and range runs on its own. Batch queries you\nhave seen come back in a few seconds. Run the expensive ones at\nmost two at a time against one datasource, and one at a time once one of them has timed out.\n\nDifferent phrases over the same selector and range are the same scan repeated. Ask for them\nin one query where the language allows it — the query-language reference shows how — rather\nthan one query per phrase.\n\n## Writing a query\n\n`log_query`, `metric_query` and `span_query` take `expression`: a query in the\ndatasource's own language, run exactly as written. A failed or empty query is yours to fix\nand resubmit. The remaining query tools, and connector tools, vary — some take a query you\nwrote, others take a plain-English description. Read each tool's description for its contract.\n\nBefore your first query against a datasource, read every file its `data-sources.yaml`\nentry lists under `docs`, as [Docs](#docs) says: the query-language reference under\n`/telemetry/references/`, and the datasource's own guidance under `/telemetry/guidance/` —\nits real labels, fields, metric names and worked examples. What a query costs and how to\nsize its range are further down each file, so read it to the end.\n\nUse only names you can source from the guidance, a result, or the incident: a guessed label or metric matches nothing, and nothing warns you that it\ncould not have matched. When none of those covers the label, label value or metric name\nyou need, read it off the datasource with `telemetry_inspect` rather than guessing.\n\nYou own query cost, and the datasource will time out or refuse expensive scans. Anchor\nevery filter to real values; a match-anything selector scans the whole estate. In many query\nlanguages only some parts of a query decide how much the source reads — an indexed selector\nand the range — while filters applied after reading make it no cheaper. The query-language\nreference says which parts those are. When a query times out, narrow those parts rather than\nfanning out variants of the same expensive scan.\n\nSet `purpose` to what the query is trying to establish — it is recorded with the query and\nshown to the responder beside your expression.\n\nTimes are UTC. [Times](#times) says how to give one a person expressed on a wall clock.\n\nEvery result carries the expression it executed, so changing one filter on a query that\nworked means copying that expression and editing it, not rewriting it from memory.\n\nWhen the question is who or which — which job wrote these rows, which caller sent the\nfailing requests, which tenant's traffic dominated — count by the field that names the\njob, caller or tenant over the whole window rather than listing lines. Filter on what you\nalready know — the table, message or resource that was touched, and the window — not on\nthe values you expect the answer to take. A filter built from expected answers can only\nreturn those answers.\n\n## When a query comes back empty\n\nEmpty is a result, not a failure, and two very different things produce it: the data\ngenuinely shows nothing, or your query matched nothing it could have matched.\n\nDo not guess between them. An empty result may carry a `diagnostics` block saying what the\ndatasource could establish about which it was — read that first, and follow what the\ntool's description says it means. Where there is no diagnostic to read, widen the query: drop\nthe narrowest filter, or stretch the window, and see whether rows appear. Rows on the wider\nquery mean your filter was wrong.\n\nWhen you report an absence, say what you established and how. \"No matching errors in the\nlast hour on this datasource\" is a finding. \"There were no errors\" is a claim you have not\nearned.\n\n## When a source refuses\n\nA refusal is the other kind of non-answer. The source declined to run your query rather\nthan running it and finding nothing, so what it tells you is about the query: it asked for\nmore than the source will scan, grouped into more series than the source will hold, or\nreached back past what the source still keeps. Re-running the same shape is the one\nresponse that cannot work.\n\nStep down instead of dropping the question, and step down on the axis the refusal names:\n\n- **Too many series** (\"maximum number of series\", \"too many series\", a cardinality\n  limit): the grouping is too fine. Drop the grouping field with the most distinct values\n  first — IDs such as organisation or user before names, names before the field that says\n  who did the work, such as subscriber, job or pool — and keep the field that answers the\n  question. Add the ID back only once you know which of those matter.\n- **Too much scanned** (\"timed out\", \"deadline exceeded\"): shorten the window, or narrow the\n  part of the query that decides what the source reads (see the reference). A timeout inside\n  a batch may be the batch: re-run that query on its own before concluding the shape is too\n  broad. If the shorter window also times out on its own,\n  the window was not the problem: change the filter or aggregate by a field instead.\n- **Capped listing** (\"Result truncated\", a line or entry cap): the source ran the query and\n  returned only the newest lines. A thousand-line cap over thirty minutes can be the last\n  few seconds. Nothing before the first returned line has been read. If the moment you want\n  is earlier, end the window at that moment and shrink it until the result fits, or count\n  by a field instead of listing.\n- **Rate limited** (\"too many outstanding requests\"): your own batch is keeping the source\n  busy. Re-run only the queries that were refused, at most two at a time. Re-running the\n  whole batch after a pause meets the same limit.\n- **Past retention**: another datasource may hold the same signal for longer;\n  `data-sources.yaml` says which, and an identifier you already have carries the question\n  across to it.\n\nNever narrow what you are asking about to recover from a refusal. A filter drawn from the\nvalues you expect the answer to take — a particular event, job or caller — makes the query\ncheaper by excluding everything else, and if the answer lies outside those values the empty\nresult then looks like an answer. Narrow the window, the stream or the grouping instead.\nNote the refusal on your way past: a retention wall does not move, and one run should only\nmeet it once.\n\n## Reporting\n\nSay what the data shows and cite the queries that showed it. Where it cannot answer the\nquestion, say that plainly — a confident wrong answer costs an engineer more than an\nhonest gap, because they will stop looking.\n\nA query that failed searched nothing. When you split a question across windows or streams\nand some of them timed out, were refused or did not parse, do not report a total or a zero\nfor the whole: give the figure for what did run, and name the windows or streams it does not\ncover. If you searched fewer streams than the question asked about, say which ones; a zero\nthere is not a zero for the question.\n\nDescribe what you searched from what the queries ran, not from what you meant to run: read\neach result's time range and selector back before you name them. Never give the requested\nrange as the one you searched when only part of it ran, and never attribute a record to an\norganisation, service or caller its own fields don't name.\n\nBefore you conclude, check that the result holds the value your conclusion turns on. A\nsummary across a family of series does not give you any one series' number, and a total\ndoes not give you the split inside it. Where the deciding value is missing, say the result\ncannot tell you. Do not read it as a yes or a no, and do not let it outweigh a more\nprecise measurement you already hold. Query the one series you need, or report the gap as\na gap.\n"
}

SHA-256 of public snapshot: da6480f68d29347380ebf2ffdc0522d826010bcfdb26fb5fef9d5bfae98e4458