← Plugin catalog
Developer Tools

incident.io

incident.io v1.20261009.616

Publisher description

From the marketplace listing

Makes your agent fluent in incident.io: responding to and investigating incidents, working with on-call schedules and escalations, and authoring the operational content — runbooks, skills, and plugins — that incident.io investigations draw on. Bundles the official incident.io MCP server.

Language: English · Automatically detected from descriptions.

Publisher keywords

Search terms declared by the publisher.

Matches for “UN”

Exact text from the indicated source. A mention alone does not establish support for your task.

Publisher keywords · listing

incident-response on-call investigations runbooks sre devops

Publisher description

Makes your agent fluent in incident.io: responding to and investigating incidents, working with on-call schedules and escalations, and authoring the operational content — runbooks, skills, and plugins — that incident.io investigations draw on. Bundles the official incident.io MCP server.

Changes

incident.io

Oct 9, 2026 · 6 saved observations

Capabilities & instructions

Instruction wording changed from “Understand the extensions so that you can help a user configure and manage” to “Set up, explain and troubleshoot incident.io extensions: the plugins of skills and the”. 17 additional added or edited lines are in the evidence.

Skill evidence →
Capabilities & instructions

Declared skills changed from “[{"description":"Answer questions about how the organisation builds, deploys, and runs its software — what a system is, where it runs, what it depends on, and the real names of things (cloud projects, clusters, namespaces, hostnames, buc...” to “[{"description":"Answer questions about how the organisation builds, deploys, and runs its software — what a system is, where it runs, what it depends on, and the real names of things (cloud projects, clusters, namespaces, hostnames, buc...”.

Metadata evidence →Listing evidence →
Technical updates

Package contents changed in 8 files: .claude-plugin/plugin.json, .codex-plugin/plugin.json, README.md, …. Open the file diff to inspect the edits.

Files evidence →
1 more changes that day

Discoverability changed from “UNLISTED” to “LISTED”.

Metadata evidence →Listing evidence →

Files & skills

File archives

Plugin package48 files · 105 KBBrowse files →
Skill instructions
architecture4.72 KB

View saved version →

---
name: architecture
description: >
  Answer questions about how the organisation builds, deploys, and runs its software —
  what a system is, where it runs, what it depends on, and the real names of things
  (cloud projects, clusters, namespaces, hostnames, buckets) — from architecture docs
  wherever they live. Use when asked "how does X run", "what is Y", "where does Z live",
  or when grounding a component before debugging it.
argument-hint: "<an estate question to answer>"
---

# Architecture

Architecture docs describe what systems *are*: where they run, what they depend on, and
the real names of things. They pair with runbooks — runbooks own *procedures* (how to
diagnose and fix a failure), architecture owns *facts* (what the component is in the
first place). This skill answers estate questions (the estate: everything you run and
where) from those docs, cited — never from general knowledge. General knowledge is
exactly what these docs exist to override: a team's setup differs from defaults in
precisely the ways worth writing down.

## Where you are running

You are a coding agent with the incident.io plugin installed.

- **Before you start:** load the `extensions` skill and have it map the estate — which
  plugins are registered, where each lives, and their sync state. Skipping it doesn't
  fail loudly. It just means you searched the local half of the estate and reported it
  as the whole.
- **Where to search:** the four places in
  [where-docs-live.md](references/where-docs-live.md), in its order. Read it before
  searching, even when you think you know where the docs are.
- **Where live config lives:** the workspace.
- **When the docs have a gap:** if the user wants it filled now, that's the
  `architecture-author` skill.
- **When the question is really "how do I fix this failure":** hand over to the
  `runbooks` skill. Its Find job owns routing a symptom to its runbook.
- **How to reply:** in the voice the `talking-to-the-user` skill sets — lead with the
  answer, be concise, one next step where there is one.

## 1. Pin the subject

Reduce the question to the system or identifier it is about: a service name, a
hostname, a cluster, a bucket, a deployment. Keep both the literal identifier (for
keyword search and grep) and the question phrasing (for semantic search).

## 2. Search

Search the places above, in their order. Start each corpus at its README — a
well-formed corpus has a "Where do I look?" routing table that resolves most questions
in one hop. Do not grep the tree before trying the map; the map exists so one hop finds
the owning file. Grep only when the map misses. Document search results carry a source
provider and generated tags; architecture-shaped documents describe systems and
infrastructure rather than procedures.

## 3. Answer from the owning doc

- **Quote identifiers verbatim** — project IDs, cluster and pool names, hostnames,
  subscription names. A paraphrased identifier is worse than none.
- **Cite the doc** each fact came from, and the authoritative config repo where the
  doc names one.
- **Respect what the docs deliberately do not hold.** Values that churn — replica
  counts, resource limits, machine types, current flag state — are pointed at, not
  copied. Answer with where the current value lives, not a number the docs never
  promised. If the live config above holds that value, read it and say where it came
  from.
- **Keep it short.** Most questions resolve to a few sentences and one or two doc
  references.

## 4. When the docs do not cover it

First make sure that is what happened. A place you couldn't reach is not a place with no
docs, and reporting an unreachable corpus as a missing one sends someone to write a
document that already exists. Name what you couldn't search.

Once it's genuinely a gap, say so explicitly. Answer from other evidence when you have
it — deploy manifests, service definitions, config, a connected system's own listing —
clearly labelled with where it came from, not the docs. Never silently substitute general
knowledge for a missing doc.

Then note the gap as a curation candidate: a question the docs could not answer is a
section waiting to be written. Handle it as "Where you are running" says.

## Rules

- Read-only, always: this skill explains; it never mutates, flips flags, or runs
  commands that change state.
- Route, don't absorb: if the question is really "how do I fix this failure", ground
  the component here, then hand over as "Where you are running" says.

## What this skill is not for

Writing or maintaining architecture docs, diagnosis and fixes (the runbook that owns the
failure), current runtime state (replica counts, flag values — the docs point at where
those live), and product or code-level documentation (API references, user guides).

Referenced files: 1

architecture-author3.24 KB

View saved version →

---
name: architecture-author
description: >
  Write and maintain architecture docs — the documents that say what each system is,
  where it runs, what it depends on, and the real names of things (cloud projects,
  clusters, namespaces, hostnames, buckets). The heart of it is an interview that pins
  down what each system name actually means before anything is written. Use when asked
  to write, extend or improve architecture documentation, to document a system, or to
  fill a gap the `architecture` skill could not answer.
argument-hint: "<the system or systems to document>"
---

# Architecture author

Architecture docs describe what systems *are*: where they run, what they depend on, and
the real names of things. They pair with runbooks — runbooks own *procedures* (how to
diagnose and fix a failure), architecture owns *facts* (what the component is in the
first place) — and each side chains to the other rather than absorbing it. This skill
writes and maintains those docs. Answering questions from them is the `architecture`
skill's job, and every write starts there: you search what already exists before
writing anything.

## Before you start

Do both of these before anything else here:

1. **Load the `extensions` skill and have it map the estate**: which plugins are
   registered, where each lives, and their sync state.
2. **Load the `architecture` skill and follow it through its search** for each system
   name in scope. It searches every place architecture docs can live. What it finds
   shapes the whole job: an existing doc means extending it, not writing a sibling.

Skipping them doesn't fail loudly. It just means you wrote a second copy of a doc that
already existed, somewhere nobody looked.

## The job

Author or extend architecture docs. The heart of it is an interview that resolves what
system names actually mean before anything is written: the names people use are
ambiguous, and boundaries are decisions the owner makes, not facts an agent infers.
→ [references/write.md](references/write.md)

Where the new docs go, and how to check they will be found, is
[references/homes.md](references/homes.md).

## The taxonomy

Architecture docs work when they follow a small structural spec — systems are
directories (one per thing responders reason about separately, regardless of repo
layout), views are root files answering one cross-system question, estate services
(observability, the data platform, CI) are directories whose README routes across
their tools, the README is the map, and churny values are pointed at rather than
copied. The spec lives in [references/format.md](references/format.md); a corpus may
carry its own FORMAT.md, which takes precedence.
[references/concerns.md](references/concerns.md) catalogs the recurring concerns
(deployment, database, events, …) and the questions each file answers, and
[references/examples/](references/examples/README.md) is a complete worked example
corpus to calibrate depth against.

## What this skill is not for

Answering questions from existing docs (that's the `architecture` skill), diagnosis and
fixes (the runbook that owns the failure), current runtime state (replica counts, flag
values — the docs point at where those live), and product or code-level documentation
(API references, user guides).

Referenced files: 14

doctor4.73 KB

View saved version →

---
name: doctor
description: >
  Review the health of your incident.io agent estate — plugins, skills, connections —
  and say what to fix, without fixing anything itself. Use whenever anyone asks
  whether their incident.io setup is healthy, what's degraded, or what to improve, in
  any phrasing — one-off or as a recurring review. Every finding routes to the thing
  that fixes it.
argument-hint: "<a plugin or area to focus on — or nothing, to be shown what can be reviewed from here>"
---

# Doctor

A review, not a repair: doctor reads the estate, ranks what it finds, and hands each
finding to the thing that fixes it. It never edits a plugin's tree (skills, README,
docs — any file), never re-syncs, never changes a setting, and never writes into the
estate — the report itself is handed over, not filed. All of this holds when it runs
on a schedule. That's what makes it safe to run weekly and boring to read when
everything is fine, which is the goal.

## The shape of a run

1. **Orient** — a fast pass, cheap reads only: where the session is, whether a plugin
   lives here and corresponds to one the account knows, what the account's own health
   rollup says, and whether this agent's own install is current. Ends in a compact
   situating block the user sees within seconds.
   → [references/orient.md](references/orient.md)
2. **Offer** — an interactive, unscoped invocation pauses on that block plus a short
   menu of the reviews that make sense from here, each with an honest cost signal,
   and waits for the pick. A scoped invocation skips the menu but still shows the
   block; a scheduled run, with nobody there to read it, skips both. The skip rules
   are in orient.md.
3. **Review** — the chosen legs, below, with a progress stream reporting each step as it
   completes so the run is legible while it runs.
   → [references/progress.md](references/progress.md)
4. **Report** — headlines name the main finding and summarize the rest. An optional
   situating block adds context; short findings carry routes, and briefs carry the
   detail. Follow [references/report.md](references/report.md).

## The legs

- **Extensions** — plugins and their skills, checked end to end: sync state, how
  skills perform once loaded, the feedback issues worth acting on, feedback
  orphaned by renames, and convention drift against the authoring rules.
  → [references/extensions.md](references/extensions.md)
- **Connections and telemetry** — what this session can and cannot read about the
  rest of the estate, stated honestly.
  → [references/telemetry.md](references/telemetry.md)
- **Content drift** — the copies that were right when written: runbooks restating what
  a skill owns, procedures naming levers that no longer exist, pointers to homes that
  don't hold the content, and skill-shaped knowledge with no skill to live in.
  → [references/content-drift.md](references/content-drift.md)

Run all three for a whole-estate review. A named plugin gets the Extensions and
Content drift legs; Connections and telemetry contributes only its orient snapshot
unless the user asks for that leg. For any other scope, run the legs the user named.

## Ground rules

- **Propose-only, always.** Every finding becomes a route: an improvement brief for
  the `skill-authoring` skill's improve job, a structural gap for the `extensions` skill, a
  change in the plugin's repository, or a named dashboard page. Doctor stops at the
  hand-off — the report is the run's last act, and acting on a brief is a new job the
  user starts, never a continuation of this one.
- **Verify before alarming.** A computed status is a hint, not a verdict: check an
  issue's anchor against the current tree, and apply the reading corrections from the
  `skill-authoring` skill's improve reference before reporting an issue as real or
  resolved.
- **An unreadable surface is reported, not guessed.** Where this session has no way
  to read part of the estate, the report says so and names where the answer lives.
  Never pad a leg with speculation to look thorough.
- **Healthy is a finding.** "All plugins synced, no actionable issues" is a complete
  and useful report line, not a failure to find something.
- **The run is visible while it runs.** Review reports each step as it finishes, per
  [references/progress.md](references/progress.md), so a review that takes minutes is
  something the user can follow rather than wait out. Those lines are status; the
  report restates everything and is what gets filed.

## What this skill is not for

Setting the estate up or growing it — that's `extensions` (which routes here for
reviews, as this skill routes there for gaps). Making the edits — that's
`skill-authoring`. And incident response: doctor reviews the machinery agents use, not
live incidents.

Referenced files: 6

extensions7.86 KB

View saved version →

---
name: extensions
description: >
  Set up, explain and troubleshoot incident.io extensions: the plugins of skills and the
  connectors (MCP servers, HTTP APIs) that give incident.io's agents an organisation's
  own knowledge and tools. Use whenever you're working with incident.io plugins, skills,
  connectors or the extension_ tools in any way, including why one isn't working: a
  skill that never loads, a change that wasn't picked up, a sync error, a connector tool
  investigations won't call, whether a skill is used or helping. Not for questions about
  the organisation's own systems, which the architecture skill answers.
---

# Extensions

Before anything else, check the model this session runs on. These workflows need
Opus/Sol level or higher; smaller models follow them unreliably. If this session is
below that level, tell the user and recommend switching model before continuing.

Extensions are how an organization gives incident.io's agents its own tools and
instructions. A **skill** is a short set of instructions an agent follows for one job.
A **plugin** is the folder in the organization's repository that holds its skills,
which incident.io reads. A **connector** is an external tool (an MCP server) agents
can call on the organization's behalf. Use these sentences when a user first meets
each term.
[references/extensions-product.md](references/extensions-product.md) explains how the
system works; read it before answering questions about it or changing anything.

This skill is the entrypoint. It teaches the system, reads the current state of an
estate, and picks the right mechanism for a problem — and it routes every specialised
job to the skill that owns it rather than doing it here.

Whatever the task, the first two moves are the same: read the estate, and hear what
the user needs. Nothing else — no drafting, no probing of any connected system —
starts before both.

## Route by the task

- **Understand the system** — what plugins, skills, and connectors are, how agents use
  them, what's possible →
  [references/extensions-product.md](references/extensions-product.md)
- **Read the estate, or set it up** — what exists today (plugins and their sync state,
  connections and their health, content), measured against what a ready estate has.
  One walk serves every starting point, from scratch to a readiness pre-check.
  → [references/estate.md](references/estate.md)
- **Diagnose why something isn't working** — a skill that never loads, a change not
  picked up, a sync error, a moved plugin, a connector tool investigations won't call, a
  trigger that never fires. The causes and the order to check them are in
  [references/extensions-product.md](references/extensions-product.md); read the
  relevant section before answering, and name the first link that fails.
- **Choose the right mechanism** — the user describes a problem ("I want these steps
  run at the start of an investigation", "I want something that diagnoses this error
  code") and it needs mapping to the feature that solves it: a skill, a runbook, an
  architecture doc, a connector, or a combination. This picks the mechanism, not a
  detailed plan of what to build — drafting the content is the owning skill's job.
  → [references/choosing-a-mechanism.md](references/choosing-a-mechanism.md)
- **Scaffold and register a plugin** — create the tree in the team's repository,
  register it, and verify the first sync.
  → [references/scaffold.md](references/scaffold.md)
- **Write or improve a skill** → load the `skill-authoring` skill *now*, before
  touching any target system — its create job owns the order of work (the user's
  context first, exploration after), and starting the exploration here skips the
  gates that make the skill worth writing.
- **Answer from architecture docs** → the `architecture` skill
- **Write or improve architecture docs** → the `architecture-author` skill
- **Find or follow a runbook** → the `runbooks` skill
- **Write, rehearse or maintain runbooks** → the `runbooks-author` skill
- **Review the estate's health** ("is our setup healthy? what's degraded?") → the
  `doctor` skill

## How to talk to the user

Two kinds of first message; tell them apart:

- **A firm brief** — a mechanism and a target are named ("a triage skill for checkout
  5xx", "register the plugin under `ops/agent`"). Follow it.
- **Exploring** — "we want to set this up", "what should we do first?", a system named
  with no mechanism. Take the lead: recommend one path and say why, instead of listing
  what's possible. [references/estate.md](references/estate.md)'s from-scratch entry has
  the default and how to pick it.

A question about one thing — why a skill didn't load, what a connector can do, whether
a change landed — gets a plain answer about that thing: the answer first, the one fix or
next step, a handful of sentences in all. No Progress block, no "things to know"
sections, and nothing about the rest of the setup unless it bears on the question; a
problem you noticed elsewhere is one line at the end, at most.

While you're driving a job — setting something up, writing or fixing a skill — each
reply has this shape:

```markdown
<the answer — a few sentences, in the user's words>

**Progress**
- [x] <done>
- [ ] <this reply's step> ← now
- [ ] <still to come>

**Next step:** <one action for the user — what it unblocks>
```

The list is fixed once agreed — same items, same words, same order; only the ticks
move. If the plan changes, say so and change it once.

Progress starts on the reply that proposes the milestones (get a yes before creating
anything) and is shown, updated, on every reply after — including by a skill that takes
the job over. Before then: answer and Next step.
[references/estate.md](references/estate.md)'s milestones section has the sequence.

Use the user's words: "I tested it", not "road test" or "fresh reader"; "your setup",
not "the estate"; "added to incident.io", not "registered"; "incident.io has picked up
your changes", not "synced". Sub-agents and verification runs are your machinery:
report the result, not the mechanism. The `talking-to-the-user` skill has the full
table.

## Ground rules

- **Ground before proposing.** Grounding means two things, and target-system
  exploration is neither: the incident.io estate
  ([references/estate.md](references/estate.md) — what's registered, connected, and
  already written; `extension_plugin_list` is the first call), and the user's intent —
  what they actually need, in their words. Probing the system a skill will cover is
  part of *authoring*, owned by `skill-authoring`, and comes after the user has
  confirmed what the skill is for. However inviting a connected tool surface is,
  exploring it before that conversation is guessing with tools.
- **Confirm before creating.** Every artefact — a directory, a registration, a doc —
  is proposed with what it will contain, and created only on a yes.
- **Requirements are few; the rest is guidance.** The hard requirements are what the
  platform needs to function:
  - a repository the connected source-control integration can read
  - a registered plugin

  Everything else (architecture docs, runbooks, more skills) is a recommendation:
  explain why it produces better results, then respect the user's choice. A narrow
  use case gets a narrow setup, not the full walk's ambitions.
- **Speak the user's language, not this skill's** — "How to talk to the user" above.
- **Connections are created in the dashboard, never here.** Connecting a tool to
  incident.io is an authentication flow. Where a gap is found, link the user to the
  dashboard's Extensions page and continue with what exists.
- **Specialised work goes to the skill that owns it.** This skill owns the
  estate-level picture and the routing; the routes above name the owners.

## What this skill is not for

Incident response — this skill configures the machinery agents use, it doesn't
investigate incidents.

Referenced files: 4

on-call8.57 KB

View saved version →

---
name: on-call
description: >
  Who is on call in incident.io and how they get paged: schedules, rotas and shifts,
  overrides, cover requests, escalation paths, and pages (escalations). Use whenever
  you're working with any of these in any way, including plain reads ("who's on call
  tonight?", "when am I next on?") and paging ("page Ava", "ack my page", "has anyone
  acked?"). Load it before the first schedule, escalation or cover-request tool call:
  the tools return raw shifts and levels, and this skill says how to read them.
---

# On-call

A **schedule** holds **rotations**, each with members and **layers**: the positions
filled at the same time, like Primary and Secondary. "Primary" names a layer, not a
rotation. Shifts are computed from that configuration, so the rota isn't stored — it's rendered
for a time window. An **override** sits on top of the rotation for a window and changes
who is on call; the shifts it produces carry an `override_id`, which is how an override
is found and undone later. A **cover request** is the polite alternative: it asks the
schedule's members to volunteer, and an accepted request creates the override itself.

Only native incident.io schedules are visible. When an organisation manages on-call in
an external provider (PagerDuty, Opsgenie…), the tools say so — that means we cannot
see who is on call, never that nobody is. An organisation that doesn't run on-call in
incident.io at all has nothing here to read or change: say so plainly instead of
hunting for schedules that can't exist.

## How you reach the tools

This section is the only part that depends on where you run. The rest of the skill names
tools by their bare names and applies however you call them.

### Calling the tools

The tools are on the incident.io connection, called by the names this skill uses.

## Reading

- `schedule_show` renders a schedule's rotations and shifts over a window you choose —
  the window may reach into the past ("who was on call last Tuesday night?"), and
  `shifts_truncated: true` means narrow the window rather than summarise a partial
  picture. A shift with `uncovered: true` is a stretch nobody is on call for, so pages
  for that layer reach no one. With an `override_id`, an override cleared it; without
  one, the rota itself leaves the gap.
- "Am I on call?" / "when is someone next on?" is one `schedule_list` call with the
  `user` filter — `"me"` for the asker, a user ID or email for anyone else. Each result
  carries that user's in-progress and next shift; don't fetch whole schedules and
  compute it yourself.
- `cover_request_list` finds cover requests — "my open request", "Milly's request" —
  and returns the `cover_request_id` the respond and manage tools need. It defaults to
  pending requests; filter by `user` or `schedule_id` to narrow.
- "Who gets *paged*" is an escalation question, not a rota question —
  `escalation_path_show` resolves current on-call at every level, and each level's
  `schedules` names the schedules it pages. That is also how to answer "what uses this
  schedule": check the paths. A reference the tools don't show is unchecked, not absent,
  so say what you couldn't see rather than that nothing depends on it. A team's path comes
  from the team: `team_show` lists the paths it owns. Never pick a path because its name
  looks like the team's.
- When you hand someone over as the person to contact, check they can actually be
  reached: look them up with `user_list` and `include_inactive: true`, then read `state`
  and `seat_type`. Without that flag, a deactivated person just doesn't appear. Only an
  `on_call` or `on_call_responder` seat can be paged. The rota still renders a
  `responder`, a `viewer` or a deactivated person on their shifts, but nobody can page
  them, not even directly. Say so, and check whoever you name as taking over next the same
  way. A plain "who's on" question doesn't need the check.
- On a page, a target with `not_paged_reason` was never paged, whatever else the record
  shows. A target without one isn't proof it arrived: check that person's seat the same
  way before saying the page reached them.
- `escalation_create` on a path returns a draft (`suggestion`), not a page. With
  `card_posted: false`, nobody sees it until you act: when the user asked for the page,
  send it with `action: execute` and only its `suggestion_id`, then say who it reaches.
  With `card_posted: true`, the card is in the conversation for the user to accept, so
  point them to it. Never call a draft paged.
- For who is on call for a service or component, resolve it through the catalog first:
  that walk ends at the right escalation path. A rota is often named for the thing it
  covers, so when the catalog has no answer, look for a schedule named for X before
  saying you cannot tell.

## Changing cover

An **instruction** changes the rota now; a **request** asks people first. "Put me on",
"take Sarah off", "cover me for the next 2 hours" (said as a decision) are overrides.
"Can someone cover my shift?", "ask Alex to take Friday" are cover requests — even when
a person is named, asking is still asking.

Overrides:

- An override changes who gets paged, immediately and for real. When the user has asked
  for it and you can tell who covers, which rota, and from when to when, create it and
  report exactly that — don't ask them to confirm first. Ask first only when the person
  or the window is ambiguous, or for a swap between unlike shifts (below).
  Reminders, cancelling the user's own pending request, and reads never need asking.
- `NOBODY` clears the layer it's placed on for the window; use it only when asked to leave
  the rota uncovered. Afterwards, say plainly that pages for that layer reach no one in
  that window, not just that nobody is on call, and name anyone still on call on the
  schedule's other layers.
- A swap is two overrides, one on each shift. When the two shifts are on different
  layers or differ in length, confirm first, and say what each person ends up with (for
  example, back-to-back weeks). Otherwise, write both and report both.
- A schedule with several rotations or layers needs `rotation_id` and `layer_id`, or
  the create is refused. `schedule_show` lists each rotation's layers by name, and
  every shift carries `layer_name`, so "primary" is the layer named Primary. Put the
  override on the layer held by the person being replaced, even if that displaces an
  existing override.
- Creating an override replaces any existing override it overlaps on the same rotation and
  layer. The create result lists these as `displaced_overrides` — tell the user who you
  displaced, and offer to narrow the window if that wasn't intended.
- To undo one, `schedule_override_delete` takes the `override_id` from the create's
  result or from the shift it produced in `schedule_show`. Deleting an override doesn't
  bring back the ones the create displaced, so when you revert such a create, recreate them
  from its `displaced_overrides` rather than calling it restored on the delete alone.

Cover requests:

- Before raising one, or suggesting who to ask, check `cover_request_list` for a pending
  request on that shift. When one exists, say who it has already asked and who has
  responded, and nudge it rather than raising another.
  `cover_request_create` refuses a duplicate anyway.
- `cover_request_create` raises one. The requester must be on call during the
  requested window — you can only ask for cover of your own shift. When the user
  names who should cover ("can Alex take my shift?"), pass just that person in
  `candidate_user_ids` — naming a person doesn't turn the ask into an override.
- Write the request `message` in the first person, as the requester.
- Candidates respond with `cover_request_respond`: accept (the override is created
  automatically), decline, or offer part of the window. The requester drives theirs
  with `cover_request_manage`: cancel while pending, accept a partial offer
  (`candidate_user_id` from the request's candidates), or nudge non-responders.
- Both need a `cover_request_id`: when the user points at a request rather than
  handing you one ("remind them", "I'll take Milly's shift"), find it with
  `cover_request_list` first.

## Answering

- Refer to schedules, overrides, and cover requests by name in the reply — never by
  ULID. The IDs a follow-up action needs are already in this
  conversation's tool results; read them from there rather than asking the user.
- State times with an explicit timezone label, rendered for the user rather than in
  raw UTC when their timezone is known.
- After a write, own what changed: who is now on call instead of whom, and until when —
  never a bare "done".
skill-authoring5.29 KB

View saved version →

---
name: skill-authoring
description: >
  Create and improve the skills in your own plugins — the ones incident.io's agents
  and your coding agents load. Use whenever you're writing, editing, or reviewing a
  skill in any way: creating one, improving one from usage feedback or a review
  brief, or asking what makes a good skill.
argument-hint: "<what the skill should do — or the skill to improve and why>"
---

# Skill authoring

A skill is instructions your agents follow: a SKILL.md with a triggering description,
plus references it loads on demand. It lives in a plugin — a repository your
organization syncs into incident.io and can also install into coding agents. This
skill owns the authoring craft — what makes a skill get selected, followed, and
helpful — and the loop that improves a shipped skill from real usage: incident.io
records each load (one agent using the skill once) and assesses it retrospectively,
and that feedback drives the improve job.

## The two jobs

- **Create** — author a new skill into the user's plugin: pin it down with concrete
  trigger examples, check nothing owns the job already, draft to the format, register
  it, ship it. → [references/create.md](references/create.md)
- **Improve** — edit an existing skill from evidence: read its usage feedback with
  improve.md's reading corrections applied, fix by theme without undoing credited
  strengths, verify the fix against the recorded issue where the session can, ship,
  confirm. → [references/improve.md](references/improve.md)

Both jobs load the same two files before drafting:
[references/format.md](references/format.md) — the structural rules (anatomy,
descriptions as triggers, environment-neutral tool and datasource naming,
registration) — and [references/what-works.md](references/what-works.md) — what the
best-performing skills share, from incident.io's own estate and the assessment of
real usage. Format answers "is this valid"; what-works answers "will this get
followed". [references/example.md](references/example.md) shows both applied to one
worked skill — read it when you want the target picture rather than more rules.

One further reference is conditional:
[references/triage-skills.md](references/triage-skills.md) — read it when the skill
should be reached for by incident.io's investigations, which use installed skills that
own an incident's system or signature in its opening minutes. Authoring for an automated
caller changes three things: the name and description do all of the selecting, the procedure
runs unattended with nobody to ask, and its findings become someone else's evidence
rather than a report to a person.

## Where the skills live

The user's skills belong in their own plugin — a repository their organization
connects to incident.io — never in this one. This plugin carries the authoring
judgment; their content stays theirs. When the user has no plugin yet, plugin setup
is the `extensions` skill's job — route there, and come back to create the first skill.

Editing happens wherever the session is: in a checkout of the plugin's repository (the
common case — changes ride the team's review flow), or by handing the user finished
files when the session can't reach the repository. Never write into a plugin the user
didn't point you at.

## What this skill is not for

Runbooks and architecture docs — those are content with their own formats, owned by
the `runbooks`, `runbooks-author`, `architecture` and `architecture-author` skills in
this plugin; a skill is instructions for an agent, not knowledge for a person. And reviewing all your
plugins at once ("are our skills healthy?") is the `doctor` skill's job — this skill
works one skill at a time, from evidence about that skill.

## Ground rules
- **Speak the user's language, not this skill's.** The user may not have read this
  file; say what you're doing plainly. While you're driving a job, each reply has this
  shape — a one-off question gets a plain answer:

  ```markdown
  <the answer — a few sentences, in the user's words>

  **Progress**
  - [x] <done>
  - [ ] <this reply's step> ← now
  - [ ] <still to come>

  **Next step:** <one action for the user — what it unblocks>
  ```

  The list is fixed once agreed — same items, same words, same order; only the ticks
  move. If the plan changes, say so and change it once.

  A progress list agreed earlier in the session (by this skill or the one that handed
  over) is shown, updated, on every reply; where none exists yet, propose one when the
  skill's scope is agreed. Use the user's words: "I tested it", not "road test" or
  "fresh reader"; "your setup", not "the estate"; "incident.io has picked up your
  changes", not "synced". Sub-agents and verification runs are your machinery: report
  the result, not the mechanism. The `talking-to-the-user` skill has the full
  table.
- **Stamp what you write.** Every skill you create or edit carries this plugin's
  version in its frontmatter, so incident.io can tell which skills came through this
  flow. Read the version from this plugin's manifest, `plugin.json` at the plugin root
  — `${CLAUDE_PLUGIN_ROOT}` in Claude Code, otherwise two directories above this
  skill's folder — and never type it from memory. If the manifest can't be read, write
  `"unknown"` and tell the user. The format is in format.md's provenance stamp section.

Referenced files: 7

talking-to-the-user4.89 KB

View saved version →

---
name: talking-to-the-user
description: >
  How the incident.io skills speak to the person in the session: what the user hears
  and what goes in the record, ending each reply with one next step, keeping a fixed
  milestone list through a multi-step job, and using the user's words instead of the
  skills' own vocabulary. Load it before you reply to the user while running any other
  incident.io skill.
---

# Talking to the user

What's specific to this plugin about how its skills speak to the person in the session.
General style — length, tone, formatting — is the user's own agent's business and isn't
set here. The reply shape (answer, Progress, Next step) is in each skill's SKILL.md.

## Only what changes what they do next

Everything a skill learns has two audiences: the user, and the record (the report a
job files — an estate report, a review, a pull request description). The record gets
the machinery: checks run and passed, tools present or absent, tool output, sync
states, why steps run in this order, what was declined and why. The user gets only
what changes what they do next: a decision that is theirs, an action only they can
take, a blocker, anything created or changed in their repository or account. When
machinery matters to them, give its consequence — "incident.io can't read that
repository yet" — not the mechanism.

## One next step

While a skill is driving a job, each reply ends with the one thing the user does now
and what it unblocks. One step, never a list. The block is for what only the user can
do; when the next move is yours, do it in the same reply. Where creating something
needs a yes, the yes is the next step. Where a choice is theirs to make, offer the
options with your recommendation marked, and the next step is "pick one". A one-off
question, a filed report and an unattended run have no block.

## Milestones

A multi-step job lays out its milestones once the goal is agreed — in dependency
order, each with what "done" looks like — and gets a yes before anything is created.
That yes covers every artefact the list names with what it will contain; anything not
on the list is proposed separately. Show the list on every reply until the job is done
— it's how the user knows where they are — after the answer, before the Next step. Once
agreed, the list is fixed: same milestones, same words, same order on every reply, and
only the ticks move. Rewriting it each turn is disorienting. When the plan genuinely
changes, say so in the answer and change the list once: a milestone that turns out
unnecessary is struck, not silently dropped; a declined recommendation is marked
declined. A skill that picks up the job mid-way keeps updating
the list it inherited. The sequence belongs to the job's own reference — for plugins
and skills, the `extensions` skill's estate reference.

## Their words, not ours

The user may not have read any file in this plugin; never rely on it. Don't cite a
reference filename, a ground rule or a step number at them, and describe what happened
rather than how you did it. The vocabulary of this plugin's references — readers, road
tests, claims lists, the estate walk, loads and funnels — is for you. Sub-agents, fresh
sessions and verification runs are your machinery: report their result, never their
existence, and don't comment on your own process ("the road test doing its job").

> Both readers are back. First rehearsal failed on two delivery rules — the road test
> doing its job. Fixed: 310 lines in `ops/skills/dashboard-data-staleness/SKILL.md`,
> plus two `ops/README.md` rows.

says:

> I tested it twice as a newcomer would; two things failed the first time and I fixed
> them. It adds 310 lines in `ops/skills/dashboard-data-staleness/SKILL.md` and two
> rows to `ops/README.md`.

| Don't say | Say |
|---|---|
| the estate, the estate walk | your setup; "checking what you have" |
| registered | added to incident.io |
| synced, sync state, sync error | incident.io has (or hasn't) picked up your changes |
| mount name | the name it shows up under |
| road-test, rehearsal, verify, a (fresh) reader, "the readers are back" | "I tested it"; "tested it as a newcomer would" |
| delivery rules, output contract, the format | what it has to include; how it has to be laid out |
| the source-control integration | incident.io's access to your GitHub or GitLab |
| the create job, the improve job | writing the skill; fixing the skill |
| the claims list | the facts I checked |
| the interview | our conversation; "what you've told me" |
| a load, assessed loads, the funnel | a time an agent used it; how it did |
| carry, earn its place, lean on, the home of, own | include, is useful, uses, lives in, is responsible for |
| this plugin (meaning incident.io's skills) | "me", or "the incident.io skills" |

Plugin, skill, connector, runbook and architecture doc are the dashboard's own words:
keep them, and explain each once, in one sentence, when the user first meets it.
telemetry13.5 KB

View saved version →

---
name: telemetry
description: >
  Query an organization's observability data — logs, metrics, traces, profiles, dashboards,
  Kubernetes, SQL databases — with the incident.io telemetry tools. Use when answering a
  question that needs evidence from its monitoring: what errors fired, when latency moved,
  which pod restarted, what a dashboard showed, what its databases hold. Not for questions
  the incident record already answers, and not for querying incident.io's own data.
---

# Telemetry

An organization's observability data lives in datasources it connected — a Loki, a
Prometheus, a Honeycomb, a Postgres. Each holds different signals and speaks a different
query language. You reach all of them through two sets of tools: the native telemetry tools,
and the extension connector tools where there are any. The next section says how to call
them; the rest of this skill applies however you call them.

## How you reach telemetry

### Native telemetry

```
telemetry_guidance_show(path: "data-sources.yaml")
log_query(datasource_id: "…", purpose: "…", expression: "…", time_from: "…", time_to: "…")
```

The telemetry tools are on the incident.io connection. Each tool's description is the
authority on it: what it does, the arguments it takes, and how to read what it returns.

Where the session has no `telemetry_guidance_show` but has `ask_telemetry`, the telemetry tools
are not enabled for this organization: ask `ask_telemetry` the question in plain words
instead, and say so. Where it has neither, say the incident.io connection has no telemetry
tools rather than guessing an answer.

#### Running queries together

Run queries that do not depend on each other as parallel tool calls, in one turn.

### Extension connector

Connector tools are on the incident.io connection, named `<connector>__<tool>`. If there are
no such tools, there are no connectors.

### Docs

Read the datasource index, `data-sources.yaml`, and each file its `docs` fields list with
`telemetry_guidance_show`, passing each path exactly as listed. Read each one whole, in one
turn of parallel calls. If your client cuts a result short or saves it to a file, read on
from where it stopped before you query.

### Times

Put the window in `time_from` and `time_to` as RFC3339: leave them out and the query covers
only the last hour. Convert anything a person expressed as a wall-clock time from the current
date in the conversation, rather than from your own sense of now.

### Errors

Read the error message, not the error code: the message says what went wrong. An internal
error with no detail means the call failed, not that the data is empty. A product error
means the organization cannot query telemetry here — say so. A datasource that cannot be
reached is a fact about the monitoring, not an answer to the question.

## Connectors

Some sources of telemetry data are made available through a connector rather than a native
telemetry data source. Connectors are not listed in `data-sources.yaml`.

When picking where a signal lives, consider both the telemetry datasources and the
connectors, then query the one that holds it. Native telemetry wins when both hold the same
data; do not query both to confirm. Note that some vendors will be available through both
native telemetry (e.g. for logs and metrics) and a connector (e.g. for inspecting ingestion
configuration). When the native datasource for a signal errors or returns nothing, check the
connectors before reporting the signal as unavailable.

Use your judgement to decide whether a connector is appropriate to use to answer a telemetry
query - not all connectors, or tools within a connector, expose telemetry data. E.g. a
connector called `production operations` with a tool `read_service_restart_logs` would be
appropriate for a telemetry query, whereas a connector called `linear` with a tool
`fetch_project_updates` would not be.

## The loop

1. **Read `data-sources.yaml`**, then **the connectors**, to pick a datasource.
   `data-sources.yaml` lists each one with its ID, type, what it holds and the earliest data
   it still keeps — and, under `docs`, the paths of its query-language reference and its own
   guidance. Match the question to a datasource that holds that signal: a metrics source
   cannot answer a question about log text. Then check the window you need against that
   datasource's earliest timestamp, rather than learning it from a refusal later.
2. **Check `memory_recall` first**, where it is available. It returns expressions already
   proven on this account for that datasource — adapt one, or replay it exactly with
   `from_query_id`, rather than writing from scratch.
3. **Read the tool's description**, and the connector tool's if you have decided to use a
   connector, then run it.
4. **Drill into what came back.** A telemetry query returns a `result_id`, and the drill-down
   tools narrow that stored result without re-querying the source.
5. **Cite what you used.** A finding without the query behind it cannot be checked.

## Ask more than one question at a time

Queries that do not depend on each other should run together rather than one after another,
as [Running queries together](#running-queries-together) shows. An incident rarely turns on a
single query, and running them in sequence spends an engineer's time for no reason.

Batch different questions, not copies of one query you have not proven yet. A batch returns
when its slowest query finishes, so it costs whatever its worst member costs, and a shape
that is too broad — or that a source has already turned down once — fails once per copy:
slicing one unproven query across a dozen windows is a dozen ways to be told the same no.
Run it once, read what came back, and fan out from there.

A datasource serves every query from one shared budget. Queries that each read a lot — a
selector the guidance says covers most of the data, a range of hours — slow each other down
when they run together, and a batch of them can all time out where each alone would have
finished. Your first query against a selector and range runs on its own. Batch queries you
have seen come back in a few seconds. Run the expensive ones at
most two at a time against one datasource, and one at a time once one of them has timed out.

Different phrases over the same selector and range are the same scan repeated. Ask for them
in one query where the language allows it — the query-language reference shows how — rather
than one query per phrase.

## Writing a query

`log_query`, `metric_query` and `span_query` take `expression`: a query in the
datasource's own language, run exactly as written. A failed or empty query is yours to fix
and resubmit. The remaining query tools, and connector tools, vary — some take a query you
wrote, others take a plain-English description. Read each tool's description for its contract.

Before your first query against a datasource, read every file its `data-sources.yaml`
entry lists under `docs`, as [Docs](#docs) says: the query-language reference under
`/telemetry/references/`, and the datasource's own guidance under `/telemetry/guidance/` —
its real labels, fields, metric names and worked examples. What a query costs and how to
size its range are further down each file, so read it to the end.

Use only names you can source from the guidance, a result, or the incident: a guessed label or metric matches nothing, and nothing warns you that it
could not have matched. When none of those covers the label, label value or metric name
you need, read it off the datasource with `telemetry_inspect` rather than guessing.

You own query cost, and the datasource will time out or refuse expensive scans. Anchor
every filter to real values; a match-anything selector scans the whole estate. In many query
languages only some parts of a query decide how much the source reads — an indexed selector
and the range — while filters applied after reading make it no cheaper. The query-language
reference says which parts those are. When a query times out, narrow those parts rather than
fanning out variants of the same expensive scan.

Set `purpose` to what the query is trying to establish — it is recorded with the query and
shown to the responder beside your expression.

Times are UTC. [Times](#times) says how to give one a person expressed on a wall clock.

Every result carries the expression it executed, so changing one filter on a query that
worked means copying that expression and editing it, not rewriting it from memory.

When the question is who or which — which job wrote these rows, which caller sent the
failing requests, which tenant's traffic dominated — count by the field that names the
job, caller or tenant over the whole window rather than listing lines. Filter on what you
already know — the table, message or resource that was touched, and the window — not on
the values you expect the answer to take. A filter built from expected answers can only
return those answers.

## When a query comes back empty

Empty is a result, not a failure, and two very different things produce it: the data
genuinely shows nothing, or your query matched nothing it could have matched.

Do not guess between them. An empty result may carry a `diagnostics` block saying what the
datasource could establish about which it was — read that first, and follow what the
tool's description says it means. Where there is no diagnostic to read, widen the query: drop
the narrowest filter, or stretch the window, and see whether rows appear. Rows on the wider
query mean your filter was wrong.

When you report an absence, say what you established and how. "No matching errors in the
last hour on this datasource" is a finding. "There were no errors" is a claim you have not
earned.

## When a source refuses

A refusal is the other kind of non-answer. The source declined to run your query rather
than running it and finding nothing, so what it tells you is about the query: it asked for
more than the source will scan, grouped into more series than the source will hold, or
reached back past what the source still keeps. Re-running the same shape is the one
response that cannot work.

Step down instead of dropping the question, and step down on the axis the refusal names:

- **Too many series** ("maximum number of series", "too many series", a cardinality
  limit): the grouping is too fine. Drop the grouping field with the most distinct values
  first — IDs such as organisation or user before names, names before the field that says
  who did the work, such as subscriber, job or pool — and keep the field that answers the
  question. Add the ID back only once you know which of those matter.
- **Too much scanned** ("timed out", "deadline exceeded"): shorten the window, or narrow the
  part of the query that decides what the source reads (see the reference). A timeout inside
  a batch may be the batch: re-run that query on its own before concluding the shape is too
  broad. If the shorter window also times out on its own,
  the window was not the problem: change the filter or aggregate by a field instead.
- **Capped listing** ("Result truncated", a line or entry cap): the source ran the query and
  returned only the newest lines. A thousand-line cap over thirty minutes can be the last
  few seconds. Nothing before the first returned line has been read. If the moment you want
  is earlier, end the window at that moment and shrink it until the result fits, or count
  by a field instead of listing.
- **Rate limited** ("too many outstanding requests"): your own batch is keeping the source
  busy. Re-run only the queries that were refused, at most two at a time. Re-running the
  whole batch after a pause meets the same limit.
- **Past retention**: another datasource may hold the same signal for longer;
  `data-sources.yaml` says which, and an identifier you already have carries the question
  across to it.

Never narrow what you are asking about to recover from a refusal. A filter drawn from the
values you expect the answer to take — a particular event, job or caller — makes the query
cheaper by excluding everything else, and if the answer lies outside those values the empty
result then looks like an answer. Narrow the window, the stream or the grouping instead.
Note the refusal on your way past: a retention wall does not move, and one run should only
meet it once.

## Reporting

Say what the data shows and cite the queries that showed it. Where it cannot answer the
question, say that plainly — a confident wrong answer costs an engineer more than an
honest gap, because they will stop looking.

A query that failed searched nothing. When you split a question across windows or streams
and some of them timed out, were refused or did not parse, do not report a total or a zero
for the whole: give the figure for what did run, and name the windows or streams it does not
cover. If you searched fewer streams than the question asked about, say which ones; a zero
there is not a zero for the question.

Describe what you searched from what the queries ran, not from what you meant to run: read
each result's time range and selector back before you name them. Never give the requested
range as the one you searched when only part of it ran, and never attribute a record to an
organisation, service or caller its own fields don't name.

Before you conclude, check that the result holds the value your conclusion turns on. A
summary across a family of series does not give you any one series' number, and a total
does not give you the split inside it. Where the deciding value is missing, say the result
cannot tell you. Do not read it as a yes or a no, and do not let it outweigh a more
precise measurement you already hold. Query the one series you need, or report the gap as
a gap.
Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
MIT
Package author
incident.io
Keywords
See publisher keywords

Declared capabilities

  • Read
  • Write

Package observed Oct 9, 2026.

Technical details
First seen
Oct 7, 2026 · 18:00 UTC
Last seen
Oct 9, 2026 · 18:00 UTC
Latest observed change
Oct 9, 2026 · 12:28 UTC
Collection status
Collected

plugin_asdk_app_6ac3f2cb2b608191b91136b4cf737358

Download plugin data (JSON)

Before you connect incident.io

How do I connect it?

Open the publisher's marketplace listing to check current availability and follow its connection instructions. This directory does not install plugins. Check the requested access and any account requirements before connecting.

Check marketplace availability ↗

Does it require paid access?

We have not established the pricing or subscription requirements for this plugin. An absent price does not mean free access.

Compare researched pricing and access models →

How can I evaluate it?

Check the declared skills and available files, then try a small task whose result you can verify. Our archived descriptions and instructions establish publisher claims, not tested runtime quality. Review sources and coverage limits.