← Honeycomb (EU)CONTENT HISTORY

Update to Honeycomb (EU)

Snapshot Sep 30, 2026 · 23:10 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "otel-genai-instrumentation",
  "description": "Guides instrumentation of GenAI/LLM applications with OpenTelemetry for Honeycomb, including content capture and agent failure detection. Trigger phrases: \"instrument my GenAI app\", \"add tracing to LLM calls\", \"trace AI agent\", \"instrument OpenAI\", \"instrument Anthropic\", \"GenAI observability\", \"trace tool calling\", \"LLM token usage\", \"instrument embeddings\", \"trace MCP\", \"GenAI metrics\", \"instrument LangChain\", \"add GenAI spans\", \"capture prompts\", \"capture LLM responses\", \"enable GenAI content capture\", \"streaming tracing\", \"trace streaming responses\", or any request about instrumenting GenAI/LLM applications.\n",
  "included_files": [
    {
      "relative_path": "references/agent-and-tool-patterns.md",
      "size_in_bytes": 11071
    },
    {
      "relative_path": "references/auto-instrumentation-setup.md",
      "size_in_bytes": 4303
    },
    {
      "relative_path": "references/content-capture-setup.md",
      "size_in_bytes": 8100
    },
    {
      "relative_path": "references/genai-attributes-catalog.md",
      "size_in_bytes": 2103
    },
    {
      "relative_path": "references/manual-instrumentation.md",
      "size_in_bytes": 29706
    },
    {
      "relative_path": "references/mcp-instrumentation.md",
      "size_in_bytes": 5718
    },
    {
      "relative_path": "references/streaming-instrumentation.md",
      "size_in_bytes": 11410
    }
  ],
  "skill_md_contents": "---\nname: otel-genai-instrumentation\ndescription: >\n  Guides instrumentation of GenAI/LLM applications with OpenTelemetry\n  for Honeycomb, including content capture and agent failure detection.\n  Trigger phrases: \"instrument my GenAI app\", \"add tracing to LLM calls\",\n  \"trace AI agent\", \"instrument OpenAI\", \"instrument Anthropic\",\n  \"GenAI observability\", \"trace tool calling\", \"LLM token usage\",\n  \"instrument embeddings\", \"trace MCP\", \"GenAI metrics\",\n  \"instrument LangChain\", \"add GenAI spans\", \"capture prompts\",\n  \"capture LLM responses\", \"enable GenAI content capture\",\n  \"streaming tracing\", \"trace streaming responses\",\n  or any request about instrumenting GenAI/LLM applications.\nmetadata:\n  version: \"1.0.0\"\n  semconv_version: \"v1.40.0\"\n---\n\n# GenAI Instrumentation for Honeycomb\n\nInstrumenting LLM and agent applications using OTel Semantic Conventions for GenAI\n(currently v1.40.0, Development status). For conceptual foundations, see\nthe **observability-fundamentals** skill.\n\n## Base OTEL Setup (Required First)\n\n**BEFORE implementing GenAI instrumentation, ensure your base OpenTelemetry configuration is complete.**\n\nUse the **otel-instrumentation** skill to configure all standard OTEL environment variables\n(OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL,\nsignal-specific endpoints, etc.) and verify basic spans are flowing to Honeycomb.\n\nGenAI instrumentation adds GenAI-specific configuration on top of that base setup.\n\n## Critical Requirements (Non-Negotiable)\n\n**BEFORE implementing any GenAI instrumentation, complete these steps in order:**\n\n### Step 1: Ask About Content Capture (FIRST!)\n\n**Stop and ask the user this question BEFORE writing any code or configuration:**\n\n> \"Do you want to capture the actual prompts and model responses in your traces?\n>\n> **Enabling content capture:**\n> - ✅ Helps debug tool call failures, planning loops, and agent deadlocks\n> - ✅ Lets you see why the model made specific decisions\n> - ❌ Captures potentially sensitive content (user prompts, model responses)\n> - ❌ May contain PII, proprietary data, or confidential information\n>\n> **Recommended for:** debugging/development, non-sensitive data, or if you have filtering\n>\n> **Not recommended for:** production with sensitive data, PII/health/financial info\"\n\n**Record their answer** — you'll need it when configuring instrumentation.\n\n### Step 2: Enable GenAI Conventions (REQUIRED)\n\n```bash\nexport OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental\n```\n\nWithout this, GenAI spans will not be created.\n\n### Step 3: Set Required Attributes on EVERY Span (REQUIRED)\n\n- `gen_ai.operation.name` — e.g., `chat`, `execute_tool`, `invoke_agent`\n- `gen_ai.conversation.id` — same value for all spans in a conversation\n\n**Impact if missing**: Spans won't be recognized as GenAI operations and cannot be queried by session.\n\n### Step 4: Implement force_flush() (REQUIRED)\n\nGenAI apps often exit early (crash, Ctrl+C, CLI). Force flush after each top-level invocation\nto prevent silent span loss.\n\nFor OTLP configuration, environment variables, and Honeycomb authentication (including the\nsilent-rejection pitfall), see the **otel-instrumentation** skill.\n\n## Prerequisites\n\n**This skill assumes your agent application is already sending telemetry to Honeycomb.** You should have:\n- OpenTelemetry SDK installed and initialized\n- All standard OTEL environment variables configured (see **Base OTEL Setup** section above)\n- OTLP exporter configured with your Honeycomb API key\n- Basic spans flowing to Honeycomb\n\n**If you haven't set this up yet, use the otel-instrumentation skill first** for:\n- SDK setup and dependencies\n- OTEL environment variables (OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_*, etc.)\n- OTLP configuration and Honeycomb authentication\n- Verification that spans are flowing\n\nOnce base telemetry is working, return here to add GenAI-specific instrumentation.\n\n## Auto-Instrumentation (Python and Node.js)\n\nPython and Node.js have official OTel auto-instrumentation packages for GenAI providers.\nGo, Java, etc. require manual instrumentation (section below).\n\n### Python\n\n| Package | Provider | Min SDK Version |\n| :--- | :--- | :--- |\n| `opentelemetry-instrumentation-openai-v2` | OpenAI | openai >= v1.26.0 |\n| `opentelemetry-instrumentation-anthropic` | Anthropic | anthropic >= v0.16.0 |\n| `opentelemetry-instrumentation-claude-agent-sdk` | Claude Agent SDK | claude-agent-sdk >= v0.1.14 |\n| `opentelemetry-instrumentation-google-genai` | Google GenAI | google-genai >= v1.32.0 |\n| `opentelemetry-instrumentation-vertexai` | Vertex AI | google-cloud-aiplatform >= v1.64 |\n| `opentelemetry-instrumentation-langchain` | LangChain | langchain >= v0.3.21 |\n| `opentelemetry-instrumentation-openai-agents-v2` | OpenAI Agents | openai-agents >= v0.3.3 |\n| `opentelemetry-instrumentation-weaviate` | Weaviate | weaviate-client >= v3.0.0, < v5.0.0 |\n\nSetup: `pip install <package>` + `Instrumentor().instrument()` or CLI\n`opentelemetry-instrument`.\n\n### Node.js\n\n| Package | Provider | Min SDK Version |\n| :--- | :--- | :--- |\n| `@opentelemetry/instrumentation-openai` | OpenAI | openai >= 4.19.0 |\n| `@opentelemetry/instrumentation-langchain` | LangChain | langchain >= 1.0.0 (not yet published to npm) |\n\nSetup: `npm install <package>` + register via OTel Node SDK.\n\nFor per-provider install commands, upstream README links, and supported version\ndetails, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/auto-instrumentation-setup.md`.\n\n## Manual Instrumentation\n\nFor languages without auto-instrumentation (Go, Java, etc.) or when\nauto-instrumentation doesn't cover your needs.\n\nKey patterns:\n- Creating inference spans (`chat`, `text_completion`, `generate_content`)\n- Creating embedding and retrieval spans\n- Setting request attributes before the call, response/usage attributes after\n- Error handling with `error.type` and span status\n\nFor code examples in Python, Node.js, and Go, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`.\n\n## Span Flushing for GenAI Apps\n\n**Critical for GenAI applications.** The `BatchSpanProcessor` buffers spans (default\n5 s schedule delay). GenAI agent runs are long-lived but may exit before the batch\nflushes — crash, Ctrl+C, short CLI invocations — causing **silent span loss**.\n\n**Rule: force-flush after every top-level agent invocation.** Expose the span\nprocessor and call `forceFlush()` without tearing down the SDK, so subsequent\ninvocations continue producing spans.\n\n### Why `shutdown()` is wrong here\n\n`sdk.shutdown()` tears down the entire pipeline — after shutdown, no new spans are\nrecorded. For apps that run multiple agent invocations (polling loops, HTTP servers,\nCLI batch modes), you need spans to keep flowing. Use `forceFlush()` instead.\n\n### Python\n\n```python\nfrom opentelemetry.sdk.trace import TracerProvider\nfrom opentelemetry.sdk.trace.export import BatchSpanProcessor\n\nspan_processor = BatchSpanProcessor(exporter)\nprovider = TracerProvider()\nprovider.add_span_processor(span_processor)\n\nasync def flush_telemetry():\n    \"\"\"Flush pending spans without shutting down.\"\"\"\n    span_processor.force_flush()\n```\n\n### Node.js\n\n```typescript\nimport { BatchSpanProcessor } from \"@opentelemetry/sdk-trace-base\";\n\nlet spanProcessor: BatchSpanProcessor | null = null;\n\nexport function initTelemetry(): void {\n  // ... exporter setup ...\n  spanProcessor = new BatchSpanProcessor(traceExporter);\n  sdk = new NodeSDK({ spanProcessors: [spanProcessor], /* ... */ });\n  sdk.start();\n}\n\nexport async function flushTelemetry(): Promise<void> {\n  if (spanProcessor) {\n    await spanProcessor.forceFlush();\n  }\n}\n```\n\n### Go\n\n```go\nvar spanProcessor *sdktrace.BatchSpanProcessor\n\nfunc InitTelemetry() {\n    spanProcessor = sdktrace.NewBatchSpanProcessor(exporter)\n    // ... provider setup ...\n}\n\nfunc FlushTelemetry(ctx context.Context) error {\n    return spanProcessor.ForceFlush(ctx)\n}\n```\n\n### Where to call `flushTelemetry()`\n\n- **After each agent invocation** — ensures the full trace (agent + chat + tool spans)\n  is exported before moving to the next task\n- **In polling/server loops** — flush after processing each request or ticket\n- **Before `process.exit()`** — as a safety net alongside `shutdownTelemetry()`\n- **NOT inside the agent loop** — flushing per-chat-turn adds latency; flush once at\n  the outer boundary\n\nExample integration:\n```typescript\nfor (const ticket of tickets) {\n  await triageIssue(ticket);   // produces invoke_agent + chat + tool spans\n  await flushTelemetry();      // ensure spans are exported before next ticket\n}\n```\n\nFor complete code examples showing flush integration with tool-calling loops, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`.\n\n## GenAI Span Types\n\n**Span names MUST follow the pattern `\"{operation} {identifier}\"`.** The `gen_ai.operation.name`\nattribute and the span name prefix must match. For example, a span with\n`gen_ai.operation.name = \"invoke_agent\"` must be named `\"invoke_agent {agent_name}\"`,\nnot `\"mypackage.DoSomething\"`.\n\n| Operation | `gen_ai.operation.name` | SpanKind | Span Name |\n| :--- | :--- | :--- | :--- |\n| Chat/completion | `chat` | CLIENT | `chat {model}` |\n| Text completion | `text_completion` | CLIENT | `text_completion {model}` |\n| Content generation | `generate_content` | CLIENT | `generate_content {model}` |\n| Embeddings | `embeddings` | CLIENT | `embeddings {model}` |\n| RAG retrieval | `retrieval` | CLIENT | `retrieval {data_source}` |\n| Tool execution | `execute_tool` | INTERNAL | `execute_tool {tool_name}` |\n| Agent creation | `create_agent` | CLIENT | `create_agent {agent_name}` |\n| Agent invocation | `invoke_agent` | CLIENT/INTERNAL | `invoke_agent {agent_name}` |\n| Workflow step | `invoke_workflow` | INTERNAL | `invoke_workflow {workflow_name}` |\n\n### Required Attributes on All GenAI Spans\n\n**CRITICAL: Every GenAI span MUST include these two attributes. This is non-negotiable.**\n\n1. **`gen_ai.operation.name`** — Identifies the operation type (`chat`, `embeddings`, `execute_tool`, `invoke_agent`, etc.).\n   - **Without this**: The span is not recognized as a GenAI operation and will be excluded from GenAI-specific queries and visualizations in Honeycomb\n   - **Set on EVERY span**: chat, execute_tool, invoke_agent, embeddings, retrieval, etc.\n\n2. **`gen_ai.conversation.id`** — Ties operations together within a conversation or session.\n   - **Without this**: Spans cannot be queried as part of a multi-operation workflow, breaking session-level analysis\n   - **Use the SAME value** across all operations in a conversation thread (user request → agent invocation → chat calls → tool executions → responses)\n   - Generate once at the start of a conversation, propagate to all operations\n\n**When to set:** When creating the span (in the span attributes), not after.\n\n**How to propagate conversation_id:**\n- In-process: Pass as parameter or store in context\n- HTTP/A2A: Include in request payload or propagate via headers\n\n**Impact of missing these attributes:**\n- Missing `gen_ai.operation.name` → Span not recognized as GenAI operation, excluded from GenAI-specific queries and visualizations\n- Missing `gen_ai.conversation.id` → Span excluded from session queries, cannot correlate operations within a conversation, breaks multi-turn analysis\n\n**What is a conversation?**\n\nA conversation is a **customer session or user interaction**, NOT a single LLM call. One conversation contains:\n- Multiple user turns/messages\n- All LLM calls handling those turns\n- All tool executions triggered by those LLM calls\n- All agent invocations within that session\n\nSee the [OTel GenAI spec](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/#conversation-id) for the definition. Key principle: use the same conversation.id when conversation history/context is maintained across operations.\n\n**When to use the same conversation_id:**\n- All operations within a single customer session\n- All turns in a multi-turn interaction\n- All LLM calls handling those turns\n- All tool executions and agent invocations within that session\n- Multiple agents participating in the same session\n\n**Example:** User starts a support session. Over the next 10 minutes they send 5 messages. The assistant makes 15 LLM calls and executes 8 tools to handle those messages. ALL of these spans share the SAME conversation.id because they're part of one customer session.\n\n**Common mistake:** Generating a new conversation_id for each LLM call. This breaks session-level analysis. Generate conversation_id ONCE at session start, reuse for all operations until session ends.\n\nFor trace structures showing how these spans compose (tool-calling loops, multi-turn\nconversations, nested agents, workflows), see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/agent-and-tool-patterns.md`.\n\n**A2A / HTTP-based agent delegation:** When agents communicate over HTTP (A2A protocol,\nREST delegation), manually propagate both trace context (via headers) AND conversation.id\n(via payload). Client: `propagation.inject()` + include conversation.id in request body.\nServer: `propagation.extract()` + `context.with()` + extract conversation.id from payload\nand pass to all operations. See the \"A2A (Agent-to-Agent) HTTP Context Propagation\"\nsection in the reference file above.\n\n## Generating and Propagating Conversation ID\n\nGenerate conversation_id at your application's **session boundary**:\n- Chat apps: when user opens new chat/thread\n- Support systems: when customer starts session\n- CLI tools: at command invocation\n- HTTP APIs: when session/conversation is created\n- Bots: when user starts thread/DM\n\nPass the SAME conversation_id to all operations within that session — all user turns, all LLM calls handling those turns, all tool executions, all agent invocations.\n\n**Propagation methods:**\n- In-process: store in session object, pass as parameter\n- HTTP/microservices: include in request payload or header (`X-Conversation-ID`)\n- Bots: store in state (Redis, DB), retrieve using thread/DM ID\n\n## Attribute Completeness\n\n**Set all attributes for which you have data available.** The OTel GenAI semantic conventions define comprehensive attributes for each operation type — if your application has the data (model name, tokens, tool arguments, etc.), set the corresponding attribute.\n\n**Critical principle**: Don't selectively omit attributes. Incomplete instrumentation limits your ability to:\n- Identify which models and agents were involved in a trace\n- Track token usage and costs across operations\n- Debug tool call failures (missing arguments/results)\n- Understand conversation flow (missing messages)\n- Correlate agent behavior with configuration (missing request parameters)\n\nFor the full attribute definitions by operation type, see the upstream semantic conventions:\n- Model operations (chat, embeddings): https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/\n- Agent operations (invoke_agent, execute_tool): https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/\n- Local reference: `${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/genai-attributes-catalog.md`\n\n**What \"data available\" means**:\n- API response fields → set corresponding response attributes (model, tokens, finish_reasons, response_id)\n- Request parameters → set request attributes (temperature, max_tokens, top_p, etc.)\n- Agent metadata → set agent attributes (name, id, description, version)\n- Tool execution → set tool attributes (name, call_id, arguments, result)\n- Conversation context → set conversation_id on ALL GenAI spans (required, not optional) — use the same ID across all operations in a conversation thread\n\nThe code examples in this skill show core attributes for each operation type. For complete coverage, consult the upstream spec and instrument every attribute your application can populate.\n\n**Impact of incomplete instrumentation**:\n\n- Missing `gen_ai.operation.name` → span not recognized as GenAI operation, excluded from GenAI queries\n- Missing `gen_ai.conversation.id` → span excluded from session queries, cannot correlate operations within a conversation\n- Missing `gen_ai.request.model` / `gen_ai.response.model` → can't identify which model was used\n- Missing `gen_ai.usage.*` tokens → can't track costs or identify expensive operations\n- Missing `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` → can't debug why tools failed or returned unexpected results\n- Missing `gen_ai.input.messages` / `gen_ai.output.messages` → can't see what prompted a response, can't debug planning loops or hallucinations\n- Missing agent attributes → can't distinguish between agents in multi-agent systems\n- Missing request parameters → can't correlate behavior with temperature, top_p, etc.\n\n**Best practice**: Instrument completely from the start. Adding attributes later requires code changes, redeployment, and waiting for new traces to arrive.\n\n## Telemetry by Failure Mode\n\nFor each failure mode, the listed telemetry enables effective debugging. Items marked\n**[Content Capture]** require enabling content capture — ask the user before enabling these.\n\n### Tool Call Failures\n\n- **Span** `execute_tool`: `gen_ai.tool.name`, `gen_ai.tool.call.id`,\n  `gen_ai.agent.name`, `gen_ai.conversation.id`, `error.type`,\n  `status.code=ERROR`, duration, `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result`\n- **Metric**: `gen_ai.client.operation.duration`\n- **[Content Capture]**: `gen_ai.input.messages` (tool_call + tool_call_response parts) —\n  shows full context of tool calls (optional, requires user consent)\n\n### Network Failures During Retrieval\n\n- **Span** `retrieval`: `gen_ai.data_source.id`, `server.address`, `server.port`,\n  `error.type`, `status.code=ERROR`, duration\n- **Metric**: `gen_ai.client.operation.duration`\n\n### Long Time-to-First-Token\n\n- **Span** `chat`: `gen_ai.request.model`, `gen_ai.usage.input_tokens`,\n  `server.address`, duration\n- **Metrics**: `gen_ai.client.operation.time_to_first_chunk` (hosted APIs) or\n  `gen_ai.server.time_to_first_token` (self-hosted)\n- Also: `gen_ai.server.time_per_output_token`, `gen_ai.agent.name`\n\n### Excessive Planning / Retry Loops\n\n- **Parent** `invoke_agent`: `gen_ai.agent.name`, `gen_ai.usage.input_tokens`, duration\n- **Children** `execute_tool`: `gen_ai.tool.name`, `gen_ai.tool.call.arguments`,\n  `gen_ai.tool.call.result`\n- **Metric**: `gen_ai.client.token.usage`\n- **[Content Capture]**: `gen_ai.output.messages` — model reasoning reveals loop cause\n  (optional but very helpful, requires user consent)\n\n### Slow Retrieval\n\n- **Span** `retrieval`: `gen_ai.data_source.id`, `server.address`, `server.port`,\n  `status.code=OK`, duration\n- **Metric**: `gen_ai.client.operation.duration`\n\n### Agent Deadlocks\n\n- **Span** `invoke_agent`: `gen_ai.agent.name`, `gen_ai.agent.id`,\n  `gen_ai.conversation.id`, `error.type=TimeoutError`, span links, duration\n- **Metric**: `gen_ai.client.operation.duration`\n- **[Content Capture]**: `gen_ai.output.messages` (tool_call parts) — reveals circular\n  delegation (optional but very helpful, requires user consent)\n\n## Content Capture (Ask User First)\n\n**CRITICAL: Do NOT enable content capture without asking the user first.**\n\n### Step 1: Ask the User\n\nBefore providing any configuration, **ask this question**:\n\n> \"Do you want to capture the actual prompts and model responses in your traces?\n>\n> **Enabling content capture:**\n> - ✅ Helps debug tool call failures, planning loops, and agent deadlocks\n> - ✅ Lets you see why the model made specific decisions\n> - ❌ Captures potentially sensitive content (user prompts, model responses)\n> - ❌ May contain PII, proprietary data, or confidential information\n>\n> Recommended if: debugging/development, non-sensitive data, or you have filtering in place\n>\n> Not recommended if: production with sensitive data, PII/health/financial info, no filtering\"\n\n### Step 2: Configure Based on Answer\n\n**If user says YES** to content capture:\n\nFor auto-instrumentation (Python), set the capture mode:\n\n```bash\n# Recommended for Honeycomb: Capture as span attributes (fully queryable)\nexport OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only\n```\n\n**Why `span_only` for Honeycomb:**\n- Content stored as span attributes → fully queryable in Honeycomb\n- Can filter, group, and visualize by message content\n- Lower overhead than `span_and_event`\n\n**Alternative modes (less common):**\n```bash\n# Events only - for high-volume scenarios where you want content in logs but not queryable\nexport OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=event_only\n\n# Both spans and events - most complete but higher overhead\nexport OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_and_event\n\n# Legacy boolean - deprecated, use span_only instead\nexport OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true\n```\n\n**Mode comparison:**\n- `span_only` → Content in span attributes (queryable, recommended for Honeycomb)\n- `event_only` → Content in events (logging, not queryable)\n- `span_and_event` → Both (most complete, 2x overhead)\n- `true` → Legacy (maps to old behavior, deprecated)\n\nFor manual instrumentation:\n- Set `gen_ai.input.messages` on chat spans (before the call)\n- Set `gen_ai.output.messages` on chat spans (after the call)\n\n**If user says NO** to content capture:\n\nDo NOT set `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` (leave unset).\n\nDo NOT include `gen_ai.input.messages` or `gen_ai.output.messages` in manual instrumentation.\n\n**ALWAYS include regardless of content capture setting:**\n- `gen_ai.tool.call.arguments` on execute_tool spans\n- `gen_ai.tool.call.result` on execute_tool spans\n\nTool arguments/results are essential for debugging and are typically less sensitive than\nfull conversation content.\n\n### What Content Capture Provides\n\nWhen enabled, `gen_ai.input.messages` and `gen_ai.output.messages` show the full\nconversation — what the user sent, what the model returned, and how tool results were\nfed back. Without them, you can see that a chat span happened but not *why* the model\nmade a particular decision.\n\n### Example: .env Configuration\n\n**If user wants content capture:**\n```bash\n# .env\n\n# Base OTEL setup - see otel-instrumentation skill for:\n#   OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT,\n#   OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL, etc.\n\n# GenAI-specific configuration (REQUIRED)\nOTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental\n\n# Content capture (OPTIONAL - ask user first)\n# Recommended for Honeycomb: span attributes (queryable)\nOTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only\n\n# Other content capture options (uncomment one if needed):\n# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=event_only  # Events only, not queryable\n# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_and_event  # Both (2x overhead)\n# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true  # Legacy (deprecated)\n```\n\n**If user does NOT want content capture:**\n```bash\n# .env\n\n# Base OTEL setup - see otel-instrumentation skill for:\n#   OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT,\n#   OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL, etc.\n\n# GenAI-specific configuration (REQUIRED)\nOTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental\n# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT not set (disabled by default)\n```\n\n### What Gets Captured\n\n**Content capture enabled** (span_only, event_only, or span_and_event):\n- `gen_ai.input.messages` — Full prompts sent to model\n- `gen_ai.output.messages` — Full model responses\n- `gen_ai.system_instructions` — System prompts\n- `gen_ai.tool.definitions` — Available tools\n\n**Capture mode determines where content is stored:**\n- `span_only` → Span attributes (queryable in Honeycomb, recommended)\n- `event_only` → Event attributes (logging/archival, not queryable in Honeycomb)\n- `span_and_event` → Both locations (most complete, double storage/overhead)\n- `true` → Legacy mode (deprecated, use `span_only`)\n\n**Content capture disabled** (default):\n- Model name, tokens, finish_reasons, timing — YES (always captured)\n- Prompt/response content — NO\n- Tool arguments/results — YES (always recommended)\n\nMessage JSON schema: `role` + `parts` (text, tool_call, tool_call_response, reasoning);\n`tool_call_response` uses `response` field (not `content`) for the tool result.\n\n### Privacy Controls (If Content Capture Enabled)\n\nIf the user enables content capture, recommend these additional safeguards:\n\n- **Filtering**: Capture selectively (e.g., exclude messages with PII)\n- **Truncation**: Limit content size (e.g., first 500 chars only)\n- **Hooks**: Route to separate access-controlled storage\n- **Access control**: Restrict who can query message content in Honeycomb\n- **Environment-based**: Full content in dev/test, disabled or filtered in prod\n\nExample filtering pattern (Python):\n```python\n# Only capture if no PII detected\nif not contains_pii(message_content):\n    span.set_attribute(\"gen_ai.input.messages\", json.dumps(messages))\n```\n\nExample truncation (any language):\n```python\n# Limit to first 500 characters\ntruncated = json.dumps(messages)[:500]\nspan.set_attribute(\"gen_ai.input.messages\", truncated)\n```\n\nFor complete setup including message JSON schemas, per-provider examples, and privacy\npatterns, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/content-capture-setup.md`.\n\n## Streaming Instrumentation\n\nStreaming (SSE, chunked responses) requires dedicated metrics and span patterns.\n\nKey metrics:\n- `gen_ai.client.operation.time_to_first_chunk` — client-observed time until first\n  streamed chunk (includes network latency); use for hosted APIs\n- `gen_ai.server.time_to_first_token` — server-side TTFT (queue + prefill); use for\n  self-hosted (vLLM, TGI)\n- `gen_ai.server.time_per_output_token` — decode speed after first token\n- `gen_ai.client.operation.time_per_output_chunk` — client-observed inter-chunk time\n\nThe span covers the full stream lifetime. Set usage attributes after stream completes.\nHandle mid-stream errors by recording the error and setting span status before closing.\n\nFor streaming span lifecycle, code examples, and error handling patterns, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/streaming-instrumentation.md`.\n\n## Evaluation Events\n\n`gen_ai.evaluation.result` event captures scoring/evaluation of GenAI output.\n\n| Attribute | Requirement | Description |\n| :--- | :--- | :--- |\n| `gen_ai.evaluation.name` | Required | Evaluation name (e.g., \"relevance\", \"faithfulness\") |\n| `gen_ai.evaluation.score.value` | Recommended | Numeric score |\n| `gen_ai.evaluation.score.label` | Recommended | Categorical label (e.g., \"pass\", \"fail\") |\n| `gen_ai.evaluation.explanation` | Recommended | Why this score was given |\n| `gen_ai.response.id` | Recommended | Links evaluation to the inference it scored |\n\nUse cases: RAG relevance scoring, hallucination detection, output quality gates.\n\n## Metrics\n\n| Metric | Type | Unit | Purpose |\n| :--- | :--- | :--- | :--- |\n| `gen_ai.client.operation.duration` | Histogram | s | End-to-end latency |\n| `gen_ai.client.token.usage` | Histogram | {token} | Input/output token counts |\n| `gen_ai.client.operation.time_to_first_chunk` | Histogram | s | Streaming TTFC |\n| `gen_ai.client.operation.time_per_output_chunk` | Histogram | s | Streaming inter-chunk |\n| `gen_ai.server.request.duration` | Histogram | s | Server-side latency |\n| `gen_ai.server.time_to_first_token` | Histogram | s | Server TTFT |\n| `gen_ai.server.time_per_output_token` | Histogram | s | Server decode speed |\n| `mcp.client.operation.duration` | Histogram | s | MCP client latency |\n| `mcp.server.operation.duration` | Histogram | s | MCP server latency |\n\nFor the required `x-honeycomb-dataset` metrics header, see the **otel-instrumentation** skill.\n\n## MCP Instrumentation\n\nModel Context Protocol instrumentation uses OTel context propagation via\n`params._meta` (W3C traceparent/tracestate).\n\n- Client spans (CLIENT) for MCP calls, server spans (SERVER) for MCP handlers\n- Key attributes: `mcp.method.name`, `mcp.session.id`, `mcp.protocol.version`\n- Metrics: `mcp.client.operation.duration`, `mcp.server.operation.duration`\n\nFor context propagation details, well-known method names, and code examples, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/mcp-instrumentation.md`.\n\n## Known Gaps & Workarounds\n\n| Gap | Workaround |\n| :--- | :--- |\n| No retry/loop count attribute | Count child spans or diff `tool.call.arguments` across siblings |\n| No inter-agent dependency (in-process) | Span links + `gen_ai.conversation.id` |\n| No inter-agent dependency (HTTP/A2A) | Manual `propagation.inject()` / `extract()` — see agent-and-tool-patterns ref |\n| No retrieval sub-metrics | Custom attributes on retrieval spans |\n| `error.type` is only error signal | Custom attributes for severity/category |\n\n## Provider-Specific Notes\n\n- **Anthropic**: cache token accounting, `gen_ai.provider.name = \"anthropic\"`\n- **OpenAI**: `system_fingerprint`, service tier, `gen_ai.provider.name = \"openai\"`\n- **AWS Bedrock**: `aws.bedrock.guardrail.id`, knowledge base attributes\n- **Azure AI**: `azure.resource_provider.namespace`\n\n## Additional Resources\n\n### Reference Files\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/auto-instrumentation-setup.md`** — Python + Node.js: per-provider install, upstream README links, supported versions\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`** — Code examples in Python/Node.js/Go for all span types\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/genai-attributes-catalog.md`** — Upstream semconv links + message JSON schema gotchas\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/agent-and-tool-patterns.md`** — Trace diagrams: tool-calling loop, multi-turn, nested agents, workflow\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/mcp-instrumentation.md`** — MCP context propagation, span conventions, method names, metrics\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/streaming-instrumentation.md`** — Streaming span lifecycle, TTFT/TTFC metrics, mid-stream errors, code examples\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/content-capture-setup.md`** — Env var + manual setup, message JSON schemas, privacy controls\n\n### Cross-References\n- **BEFORE using this skill**: Use **otel-instrumentation** for base SDK setup, all OTEL environment variables (OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_*, OTEL_EXPORTER_OTLP_HEADERS, etc.), OTLP config, collector, and sampling\n- For conceptual foundations of wide events and high cardinality: **observability-fundamentals** skill\n- After instrumenting, use the **query-patterns** skill to verify GenAI data in Honeycomb\n"
}

SHA-256: 0517dd0a85a7182e6179a1a60fd677c8a909b3a39e704d35e6a049f314501954