Flowlines
Flowlines v1.0.0
Publisher description
From the marketplace listing
Connect your Flowlines account to review AI agent activity, compare outcomes across time periods, investigate unsuccessful sessions, and analyze user cohorts. Use aggregate metrics, signals, topic maps, and focused session evidence to understand findings and identify product needs. Save independently verified findings as private workspace notes when requested. Requires a Flowlines account with access to a workspace containing agent data. Authorized results can include end-user conversation and activity data; responses should summarize evidence and mask direct personal identifiers. Flowlines records tool activity, arguments, results, and conversation outcome reports in private telemetry.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Skill instructions
flowlines-cohort-builder4.81 KB
---
name: flowlines-cohort-builder
description: Turn a question about a group of users ("which users do X", "who is at risk", "how do power users differ") into a Flowlines cohort - size it, express it in Flowlines cohort rules, hand the definition to the app, and compare it against a baseline over the Flowlines MCP server. Use when the user asks about segments, populations, or comparing groups of users.
---
# Flowlines cohort builder
A cohort is a saved rule set over a user aggregate that Flowlines keeps up to date and can compare against other cohorts. This skill sizes a candidate group with the MCP tools, writes the definition in the exact rule vocabulary, and verifies it once it exists.
## Conventions for every Flowlines tool call
- Every tool takes `reason` and `user_intent`. Keep `user_intent` identical across the conversation, for example "Identify users who churned after a frustrated session".
- Start with `get_workspace` once, then `get_context`. Pinned notes may already describe the population the user is asking about.
- Cohorts are about people. Report them in aggregate; open individual users only to validate a rule, and mask identifiers.
- End with `report_outcome` as the last tool call.
## What the MCP server can and cannot do
The MCP server reads cohorts and compares them: `list_cohorts`, `get_cohort`, `compare_cohorts`, plus `list_users`, `get_user_population_map`, and `aggregate_sessions` with `cohort_ids` and `group_by: ["cohort"]`. Creating or editing a cohort happens in the Flowlines app, in the cohort builder on the Users view, or through the REST API with a signed-in user session. Namespace API keys only authenticate ingestion, so this skill produces a definition ready to paste into the builder and verifies the result afterwards.
## Steps
1. Restate the question as a rule. Every cohort rule has a `kind`; the vocabulary and operators are in [references/cohort-rules.md](references/cohort-rules.md). Examples:
- "users who had a frustrated session in the last 14 days": `facetSentiment includes ["frustrated"]` with `windowDays: 14`.
- "power users": `numeric powerScore gte <threshold>` or `numeric sessionCount gte <n>`.
- "at-risk accounts": `numeric successRate lt 50` (rates are 0 to 100) and `recency inLastDays 30`.
- "everyone who asked about refunds": `intentFamily isOneOf ["refund"]`, using intent names from `aggregate_sessions` grouped by `intent`.
2. `list_cohorts`. Reuse an existing or system cohort when one already expresses the question; note `matchedCount`, `totalCount`, and `matchedPercent`.
3. Size the candidate group before creating anything. Use `list_users` with `order_by`, `minimum_sessions`, and `range` for count-based rules, `get_user_population_map` for a distribution, and `aggregate_sessions` grouped by `user` with the relevant filter for intent- or outcome-based rules. Report the estimated size and the identified share of sessions: `aggregate_sessions` with metric `session_count`, once with `include_unidentified: true` and once with `false`, divided. A cohort over mostly unidentified sessions is not actionable.
4. Validate the rule on two or three users from the estimate with `get_user_activity`. Confirm they belong for the reason the user cares about, not by coincidence.
5. Hand over the definition: name, one-line description, the aggregate (the users aggregate unless the user says otherwise), and the rules as JSON in the vocabulary. The app's rule suggestions can propose alternatives; mention them when the rule is a stretch.
6. After the user creates it, `get_cohort` to confirm `ruleCount`, `matchedCount`, and `totalCount` match the estimate. A large mismatch usually means a `windowDays` or a threshold differs from the estimate's range.
7. Compare: `compare_cohorts` against the baseline (usually the all-users system cohort or a mirror-image cohort), with `range` `7d`, `30d`, or `all`. The comparison reports entity count, session count, clean session rate, success rate, average cost, average latency, and signal rate for both sides; the rates are percentages from 0 to 100. `get_metric_definition` before interpreting a rate.
8. Pin the definition and the headline comparison with `save_note` only if the cohort will be used again, and never with per-user detail.
9. `report_outcome`.
## Output shape
```
Question, restated as a rule
Definition: name, description, aggregate, rules (JSON)
Estimated size: users matched of total, identified share, top examples validated (masked)
Comparison vs baseline: table of the compare_cohorts metrics
What the cohort is good for and what it cannot tell you
Open questions (also sent as unmet_needs)
```
## Do not
- Do not create a cohort over unidentified users and present it as a customer segment.
- Do not report a rate from a cohort smaller than the metric's sample floor.
- Do not list users by name in the answer; counts and masked examples only.
Referenced files: 2
flowlines-investigate-session4.27 KB
--- name: flowlines-investigate-session description: Investigate a specific problem in Flowlines - a signal, a user complaint, a failed session, or a suspicious pattern - from the aggregate down to the exact turn, over the Flowlines MCP server, while handling end-user data with care. Use when the user asks why something went wrong, wants to look into a signal or a user's sessions, or needs evidence for a bug report. --- # Flowlines session investigation Go from a symptom to a verifiable cause with the smallest exposure of end-user content. Aggregates first, then the session, then only the turns that matter. ## Conventions for every Flowlines tool call - Every tool takes `reason` and `user_intent`. Keep `user_intent` identical across the conversation, for example "Find out why the billing agent fails on refund requests". - Start with `get_workspace` once, then `get_context` for the namespace. Read pinned notes first: the problem may already be a known artifact. - Flowlines data is real production conversations. Quote the minimum, mask names, emails, phone numbers, addresses, identifiers, and payment details, and never copy a transcript into a note, ticket, or report. - End with `report_outcome` as the last tool call. ## Starting points Pick the entry that matches what the user has: - **A signal.** `list_signals` to find it, then `get_signal` for evidence and affected sessions and users. - **A user.** `get_user_activity` for the timeline, outcomes, agents, intents, signals, and representative sessions. Prefer this over listing that user's sessions. - **A session id.** `get_session` directly. - **A description.** `search` with `types` narrowed to `session` or `turn`, a time window, and the agent, then `fetch` the locators worth opening. - **A pattern.** `aggregate_sessions` grouped by `intent`, `agent`, or `day` with `failure_rate` and `failure_count`, ordered by the metric, to find where the problem concentrates before opening anything. ## Steps 1. Scope the problem with aggregates. Establish how many sessions, users, and days are affected and whether it is growing. `get_metric_definition` for any rate you quote. 2. Choose at most five representative sessions: the earliest, the most recent, and the ones the signal or user activity points at. `list_sessions` with `outcome`, `user_feedback`, `intent_family`, `from`, and `to` narrows the pick. 3. `get_session` for each. Read the analysis and the turn tree before any content: the session summary, outcome, and findings often name the failing turn. 4. `get_turn` only for the turns the analysis points at. Compare the user's request, the tool calls, and the assistant reply on that turn. Look for the usual causes: wrong intent classification, a tool error surfaced as a normal reply, missing context, a policy refusal, a loop, or a truncated response. 5. Check whether the cause is upstream of the agent: an ingestion gap (turns missing, timestamps out of order), an unmapped user identity, or an evaluator artifact. Pinned notes and `list_agent_attributes` help here. 6. Confirm the cause on a second session before calling it a cause. One session is an anecdote. 7. Estimate blast radius with one more `aggregate_sessions` filtered to the intent or agent, and check `list_signals` for a signal that already tracks it. 8. If the finding is durable and verified, `save_note` with the finding, the evidence in aggregate terms, and how future analysis should account for it. Never include personal data in the note. 9. `report_outcome`. ## Output shape ``` Symptom: what was reported, in one line Scope: sessions, users, agents, first and last occurrence, trend Cause: the mechanism, with the turn-level evidence (masked, minimal quotes) Confirmed on: session ids used as evidence Blast radius: share of sessions for the intent or agent in the window Related signals and notes Recommended fix or next check Open questions (also sent as unmet_needs) ``` ## Handling content - Reference sessions and turns by id and index, not by pasting them. - When a quote is unavoidable, keep it to the fragment that proves the point and replace identifiers with placeholders such as `<email>`. - Do not reproduce prompts, system instructions, or tool payloads in full. - If the user asks for a full transcript, point them to the session in the Flowlines app rather than reproducing it here.
Referenced files: 1
flowlines-mcp-observability11.2 KB
--- name: flowlines-mcp-observability description: Integrate, repair, review, or verify Flowlines observability in an MCP server repository. Use when a server must emit canonical Flowlines MCP tool-call telemetry through AGNTCY Observe or vanilla OpenTelemetry; do not use for Claude Code or Codex CLI telemetry. --- # Flowlines MCP Observability Instrument an MCP server so complete tool executions arrive in Flowlines as canonical MCP calls. Modify the target server; do not add a Flowlines runtime SDK or assume ownership of unrelated telemetry. ## Consent and secrets Before changing code or deployment configuration: 1. Explain that supported MCP spans export validated tool arguments, final client-visible results, and user identity metadata to Flowlines. Identity metadata includes a stable user ID and, when available, name and email; these fields and payloads may contain personal data, customer data, source code, file content, or other sensitive values. 2. Obtain explicit consent for payload export. Do not infer it from a generic request to "add telemetry." 3. Confirm the user has a Flowlines namespace API key before touching the deployment. If they do not, tell them to create one in the Flowlines app under Settings, API keys, at `https://app.flowlines.ai/settings` for the namespace that should receive the data, and offer to open that page for them (`open` on macOS, `xdg-open` on Linux). The key is shown once at creation. Ask the user to place it in the target deployment's secret manager. Never request the key in chat, write it into source or examples, interpolate it into a command, or print an existing value. 4. Treat the integration request as permission to edit and test the target repository, not to deploy it, call production tools, or mutate any production database. Read [references/contract.md](references/contract.md) before implementing or reviewing an integration. ## Inspect the target first Read the repository instructions, architecture documentation, and testing strategy. Then identify: - language, runtime, MCP SDK and transport; - package manager, lockfile, dependency policies, and supported runtime versions; - the central tool-registration or dispatch boundary; - existing MCP-level middleware or interceptors; - existing OpenTelemetry provider, exporter, collector, propagation, and shutdown handling; - where validated arguments, request ID, request `_meta`, authenticated user ID/profile, final MCP result, and error mapping are available; - how deployment secrets and environment variables are declared without values. Preserve the target's package manager and telemetry ownership. Reuse an existing tracer provider and collector when present; never register a competing global provider or replace unrelated exporters. ## Choose the integration path - For a compatible Python server using the official `mcp` package, prefer AGNTCY Observe. Read [references/python-agntcy.md](references/python-agntcy.md). - For TypeScript or any other language with an OpenTelemetry SDK, use vanilla OpenTelemetry. Read [references/vanilla-opentelemetry.md](references/vanilla-opentelemetry.md). - On the vanilla path, prefer the framework's existing MCP-level middleware or interceptor at `tools/call` as the default span boundary. Typical hooks: Go `AddReceivingMiddleware`, FastMCP `on_call_tool`, official Python `server.middleware` filtered to `tools/call`, or the equivalent TypeScript hook. Do not use HTTP, transport, or sending middleware as the Flowlines MCP span boundary; resolve identity from those layers when needed, then emit the span at MCP `tools/call`. - If automatic instrumentation or that middleware hook cannot observe the final client-visible result, validated arguments, or request metadata, keep a single wrapper around the central tool execution boundary and capture the missing fields there. Do not scatter nearly identical span code across every handler unless the framework provides no shared boundary. Do not emit Flowlines MCP spans for `initialize`, `tools/list`, or other non-`tools/call` methods. Disable overlapping automatic coverage so each call produces one Flowlines MCP span. If the stack has neither supported AGNTCY instrumentation nor a usable OpenTelemetry SDK, explain the gap instead of inventing an unverified exporter or protocol adapter. ## Implement the contract Make the smallest coherent change that satisfies all of these invariants: 1. Require non-empty `reason` and `user_intent` strings in every ordinary tool input schema. Do not synthesize either value from prompts or tool arguments. Update server instructions, examples, affected callers, and tests because this is an intentional schema change. 2. Register `report_outcome` exactly as described in the contract and include its unconditional final-call instruction in the server instructions. 3. Start one server span around each complete, validated `tools/call` execution. Give every invocation a fresh tool-call ID that is independent of the JSON-RPC request ID. 4. Record the canonical attributes from `contract.md`, the validated tool-argument object, and only the final MCP result returned to the client. When the tool has a published description, emit it as `gen_ai.tool.description` from the registration metadata, trimmed and capped at 10,000 characters. Omit missing descriptions; do not infer them from arguments or reasons. Emit the tool's published input schema as `gen_ai.tool.input_schema`, and its output schema when declared as `gen_ai.tool.output_schema`, serialized whole from the same registration metadata; omit a schema that is missing or would exceed 50,000 characters. 5. Put a non-empty, stable user identifier on every emitted MCP span as the exact `user.id` attribute. Prefer a verified authenticated subject; otherwise require client `_meta["user.id"]`. Never substitute email, display name, session ID, trace ID, or OAuth client ID. If neither identity source exists, the integration is incomplete: extend the authentication or client metadata contract rather than inventing an identity. 6. When verified profile name/email exists, emit it on the same span as exact `user.name` and `user.email` attributes. Otherwise promote non-empty client metadata as untrusted analytics values and document that provenance. Verified fields always win. Flowlines does not map name or email merely because they remain nested in MCP `_meta`; treat them as PII and never put them in captured tool arguments. 7. Configure and verify the applicable Flowlines identity mapping with user ID attribute `user.id`, name field ID `name` mapped to `user.name`, and email field ID `email` mapped to `user.email`. Use the caller-agent users mapping when a real caller agent is present, or the equivalent namespace identifier mapping for an agentless MCP session. Never label the MCP server as a caller agent. Sending the attributes alone is not sufficient for name/email profile enrichment when identity fields have not been mapped; if neither mapping surface is available, report that limitation explicitly. 8. Prefer client-supplied `_meta["session.id"]`. Never derive a conversation from user identity, trace ID, timing, or a reused protocol request ID. 9. Propagate valid incoming W3C trace context when the transport exposes it. Do not make trace context a prerequisite for a call to be recorded. 10. Mark every completed call explicitly: set span status to `OK` after a successful final MCP result and `ERROR` for a tool or protocol failure. Do not leave a completed call at the OpenTelemetry default `UNSET`, because Flowlines reports that call's success as unknown. On failure, record only a bounded error type; do not record raw exceptions, stack traces, authorization headers, OAuth claims, request `_meta`, environment variables, or secret-bearing diagnostics. 11. Keep telemetry fail-open. Export failure must not change the MCP response, and shutdown flushing must be bounded. 12. Configure OTLP through environment variables or the existing collector. Commit only secret placeholders and variable names. Do not change sampling for an application-wide provider without explicit approval. A dedicated MCP provider may use always-on sampling because these spans are product facts; with a shared provider, preserve its policy and call out any risk from unsampled remote parents. ## Verify Add tests at the middleware or wrapper boundary, using the stack's in-memory exporter when available. At minimum cover: - a successful call with explicit `OK` span status, required attributes, distinct call/request IDs, session identity, stable `user.id`, arguments, and result; - the registered tool description as `gen_ai.tool.description`, with trimming and the 10,000-character bound, plus omission when no description exists; - the registered input and output schemas as `gen_ai.tool.input_schema` and `gen_ai.tool.output_schema`, serialized whole and identical to what `tools/list` publishes, plus omission when absent or over the 50,000-character bound; - exact `user.name` and `user.email` span attributes for both the verified-profile path and the client-metadata fallback when those values are available; - a failed call with explicit `ERROR` span status that exports only the safe client-visible error and a bounded error type; - absence of `_meta`, authorization material, raw exception messages, and spoofed identity; verified identity must win over all client-supplied user fields; - `report_outcome` schema and server instructions; - exporter shutdown or force-flush behavior when the integration owns the provider. Run the target repository's narrow tests, formatter/linter, type checker, and package-manager checks. Never put a real API key in a test. Only perform live verification when the user has authorized network export and configured the key outside chat. Make ten harmless calls sharing a test `session.id` and stable test `user.id`, include a test name/email when those fields are supported, then make one final `report_outcome` call. Confirm Flowlines shows eleven accepted calls, reports successful calls as successful rather than unknown, maps all calls to the expected user ID, displays the mapped name/email, and shows the session intent, outcome, captured evidence, client attribution when supplied, and no persistent ingestion-quality issues. Behavioral clustering and tool-loop signals have separate volume and timing thresholds, so do not treat their immediate absence as exporter failure. Verifying receipt and the identity mapping needs the Flowlines MCP server signed in, or the Flowlines app. If the MCP server is connected but not authorised, ask the user to sign in first (`/mcp` in Claude Code, `codex mcp login flowlines` in Codex) rather than reporting the mapping as unverified. If neither is available, name the exact mapping to configure (`user.id` as the user ID, `name` to `user.name`, `email` to `user.email`) and where in the app to do it, and say that receipt was not verified. ## Hand off Report: - files and dependencies changed; - where deployment must set the endpoint, API-key header, and service name; - the source of `user.id`, availability of name/email, and the exact Flowlines user mappings verified; - schema or client compatibility changes caused by `reason`, `user_intent`, or `report_outcome`; - checks run and whether live Flowlines receipt was verified; - any identity, propagation, sampling, payload, or shutdown limitation that remains.
Referenced files: 4
flowlines-release-check6.02 KB
--- name: flowlines-release-check description: Check whether a release of an agent changed its outcomes in Flowlines - compare sessions, success rates, intents, cost, and signals before and after a deploy over the Flowlines MCP server. Use when the user asks whether a deploy, prompt change, or new version regressed or improved anything, or wants a go/no-go after shipping. --- # Flowlines release check Answer "did the last release change anything?" with numbers that have the right denominators, evidence from real sessions, and an honest statement of what cannot be known yet. ## Conventions for every Flowlines tool call - Every tool takes `reason` and `user_intent`. Keep `user_intent` identical across the conversation, for example "Check whether the 2026-09-01 release of the support agent regressed outcomes". - Start with `get_workspace` once, then `get_context` for the namespace. Read pinned notes: a known evaluator artifact or ingestion gap changes how a comparison reads. - Quote the minimum from sessions and mask identifiers. Prefer aggregates. - End with `report_outcome` as the last tool call. ## Inputs 1. The agent. `list_agents` shows the names Flowlines observes; use the exact name in filters. 2. The release boundary as an ISO 8601 timestamp. Take it from the user, the deploy log, or the release list in the Flowlines app under Versions. Flowlines records a release either from a reported agent version or from a detected prompt change; the app's release receipt shows which. 3. The comparison window. Default to the same number of days on each side of the boundary, at least 3 and at most 14, so both sides have comparable traffic. Say when the after-window is still short. ## Steps 1. `get_changes_since` with `since` set to the release boundary. This gives the after-side activity, outcomes, signals, and pinned notes in one call. 2. Build the two windows from day buckets. `aggregate_sessions` only accepts trailing ranges that end now (`24h`, `7d`, `30d`, `90d`, `all`), so any range you pick contains post-release activity. Query the smallest range that covers both windows with `group_by: ["day"]` and metrics `session_count`, `user_count`, `success_count`, `failure_count`, `total_cost_usd`, `unanalyzed_count`, filtered by `agent_name`. Assign each day bucket to the before or after window yourself. The deployment day is mixed: exclude it from both windows unless the boundary falls at the start of the day, and say which choice you made. Day buckets cannot represent a boundary inside a day exactly; say so, and if that precision matters use `list_sessions` with `from` and `to` around the boundary for counts on that day only. 3. Compute rates from the summed counts, never by averaging per-day or per-group rates: success rate for a window is the sum of `success_count` divided by the sum of `success_count` and `failure_count` across its days. A status-grouped query is not a shortcut, because each status group has a rate of 0 or 100 percent by construction. `get_metric_definition` for `success_rate` and any other rate you report, and apply the sample floor to the window's combined denominator; below it, report counts and say the rate is not yet meaningful. Sessions in the after-window are still being analysed, so a high `unanalyzed_count` makes the after-side rate provisional. 4. Find what moved: `aggregate_sessions` for the same covering range with `group_by: ["day", "intent"]` and `session_count`, `success_count`, `failure_count`, filtered by the agent. Split the day buckets into the same two windows and compare per intent. Intents that appear or disappear between the windows are as important as rate changes, but only call an intent new or gone when both windows have enough sessions for its absence to mean something. 5. `list_signals` for the covering range. Signals that fired only after the boundary are candidate regressions; `get_signal` for evidence and affected sessions. 6. Evidence: `list_sessions` filtered by `agent_name`, `from` the boundary, and `outcome: "unsuccessful"` or `user_feedback: "negative"`. Open at most three with `get_session`, and only the turn the analysis points at with `get_turn`. Then check one comparable before-window session for the same intent, so the difference is attributable to the release and not to the intent. 7. Rule out confounders before concluding: a traffic mix shift (different intents or users after the boundary), an ingestion gap on either side, a concurrent change to another agent, and the analysis lag. Pinned notes and `list_agent_attributes` help with the first two. 8. If the app's release receipt is available, reconcile with it: it reports claim consistency, the delta versus the previous release, the production window, the prompt diff, sample sessions, and related signals. Those are fractions between 0 and 1. Report disagreements between your comparison and the receipt rather than picking one. 9. If the regression or improvement is confirmed on more than one session and the numbers clear the sample floor, `save_note` with the finding, the boundary, the metrics before and after, and the intents involved. No personal data. 10. `report_outcome`. ## Output shape ``` Agent, release boundary, windows compared (days before / after), sessions on each side Verdict: improved / regressed / unchanged / too early, with the reason in one line Outcome metrics: before vs after table, with unanalysed share and sample floor notes Intents that moved: new, gone, or shifted, with counts Signals since the release: severity, evidence, affected sessions Evidence sessions: ids and the turn each one points at (masked, minimal quotes) Confounders considered and how they were ruled out Recommendation: keep, watch, or roll back, and what to re-check tomorrow Open questions (also sent as unmet_needs) ``` ## Do not - Do not compare a 12-hour after-window with a 7-day before-window without saying the after-side is provisional. - Do not report a rate below its sample floor as a change. - Do not attribute a change to the release when the intent mix moved at the same time. - Do not paste transcripts; cite session ids and turn indexes.
Referenced files: 1
flowlines-weekly-review5.95 KB
--- name: flowlines-weekly-review description: Produce a periodic review of a Flowlines namespace over the Flowlines MCP server - what changed since the last review, which signals fired, where outcomes moved, and what to pin for next time. Use when the user asks what changed, wants a weekly or monthly review, or asks for a status report on their agents. --- # Flowlines weekly review Turn the Flowlines MCP tools into one repeatable review with a fixed shape, so successive reviews are comparable and nothing already known is re-derived. ## Conventions for every Flowlines tool call - Every tool takes `reason` (why this call) and `user_intent` (the user's goal, identical across the conversation). Set `user_intent` once, for example "Weekly review of the <namespace> namespace". - Start with `get_workspace` once, pick the namespace, then `get_context` for that namespace. Read its pinned notes and glossary before analysing: notes record known caveats and confirmed findings. - Prefer aggregates over transcripts. Open a session only as evidence for a claim, quote the minimum, and mask names, emails, and other identifiers. - End with `report_outcome` as the last tool call, before the final answer, with `unmet_needs` for anything the server could not provide. ## Inputs Ask for, or infer, three things: 1. The namespace. If `get_workspace` shows exactly one, use it silently. 2. The review window. Default to the last 7 days. `get_changes_since` needs an ISO 8601 `since` timestamp and clamps to the trailing 90 days; take `since` from the previous review's pinned note when one exists, otherwise from the window. 3. Focus agents, if the user names any. Otherwise cover every agent with activity. ## Steps 1. `get_context` with `overview_range` matching the window (`7d` for a weekly review, `30d` for a monthly one). Record the context status; if it is `not_configured`, `gathering`, or `failed`, say so and continue without it. 2. `list_notes`. Find the most recent note whose title starts with `Weekly review` and take `since` from its body. Titles carry the review date, so each period has its own note. 3. `get_changes_since` with `since`. This returns session activity and outcomes in the window, signals that fired, and notes pinned. It replaces separate list calls. 4. `aggregate_sessions` for the window with `group_by: ["agent"]` and metrics `session_count`, `user_count`, `success_count`, `failure_count`, `success_rate`, `total_cost_usd`, `unanalyzed_count`. Run it again with `group_by: ["day"]` for the trend. Never average rates across groups or days: a rate for any combination of groups is the summed `success_count` divided by the summed `success_count` plus `failure_count`. When a rate looks surprising, `get_metric_definition` for it before interpreting: denominators, missing-data rules, and sample floors differ by metric. 5. `list_signals` for the window. For each signal that is new or grew, `get_signal` for its evidence. Group signals by agent and severity. 6. For the largest movements, `aggregate_sessions` with `group_by: ["intent"]` on the affected agent to see which intents drive the change, then `list_sessions` filtered by `outcome: "unsuccessful"` or `user_feedback: "negative"` and open at most three sessions with `get_session` as evidence. 7. Check identity coverage: `aggregate_sessions` with metric `session_count`, once with `include_unidentified: true` and once with `false`. The identified share is the second count divided by the first. `user_count` cannot show this gap because it never counts empty user ids. A low identified share means identity mapping is incomplete; report it rather than drawing user-level conclusions. 8. Decide what to pin. `save_note` only for durable, verified findings that the next review must not rediscover: an evaluator artifact, an ingestion gap, a confirmed regression and its cause. Never put end-user personal data in a note. 9. Save the continuation point. Pin one note titled `Weekly review <review date>`, for example `Weekly review 2026-09-02`, whose body records the review timestamp to use as the next `since`, the window covered, and the two or three headline numbers. The date in the title makes each period's note distinct; the server rejects a title it has seen recently regardless of the body, and notes cannot be edited. If the save is rejected as a duplicate, the period was already reviewed: `list_notes`, confirm the existing note covers this window, and reuse its timestamp. Pass `allow_duplicate: true` only when the existing note is for a different window and the title collided anyway. Confirm the save succeeded and the timestamp is recorded before reporting the review as complete. 10. `report_outcome`. ## Output shape Use this structure every time so reviews line up week over week: ``` Namespace, window, review timestamp (use it as the next `since`) Headline: 3 bullets, each with the number and the change versus the previous window Outcomes by agent: table of sessions, users, success rate, cost, unanalysed Signals: new, grown, resolved - each with severity, agent, one-line evidence Notable intents: the intents behind the largest movements Data quality: unanalysed share, unidentified users, context status, anything pinned as a caveat Pinned this review: titles of notes saved, including the dated review note and its timestamp Open questions: what the server could not answer (also sent as unmet_needs) ``` Report numbers with their metric definitions in mind: an `unanalyzed_count` above a few percent makes `success_rate` provisional, and a rate below the metric's sample floor is not a trend. Say when a comparison is against a partial previous window. ## Do not - Do not page through `list_sessions` to build totals; `aggregate_sessions` exists for that. - Do not average rates across groups; recompute from summed counts. - Do not quote prompts or responses beyond the minimum needed to support a finding. - Do not pin conversation scratch, hunches, or per-user observations. - Do not skip `report_outcome`, including when the review is empty or blocked.
Referenced files: 1
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Flowlines
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 00:00 UTC
- Collection status
- Collected
plugin_asdk_app_6aa11dfaeb20819187226d4810e1d94a
Download plugin data (JSON)