{"id":7975,"plugin_id":"plugin_asdk_app_6a51b29fa2d48191b215ff28f9a64fb4","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:51:43.474Z","digest":"ed6d5ea6b71820807187bf351ad49ac9d5f8a8ef3ebfb080831a85822b8209ab","against":null,"payload":{"name":"otel-migration","description":"Guide for retrofitting OpenTelemetry into an existing, uninstrumented application. Trigger phrases: \"migrate existing app to OTel\", \"add OpenTelemetry to existing project\", \"retrofit OTel into my codebase\", \"thread context through my code\", \"context propagation\", \"bridge Prometheus metrics to OTel\", \"logging bridge\", \"migrate logging to OTel\", \"slog bridge\", \"logback bridge\", \"verify my instrumentation\", \"traces are disconnected\", \"orphaned spans\", \"migrate to OpenTelemetry\", \"OTel migration plan\", \"how do I sequence an OTel migration\", \"add tracing to existing code\", \"refactor for context propagation\", \"Fiber context gotcha\", \"keep existing logging working with OTel\", \"add OTel without breaking Prometheus\", \"bridge existing metrics\", \"coexist with existing monitoring\", or any request about retrofitting OpenTelemetry into an existing application. This skill is for migrating existing codebases, NOT greenfield instrumentation (use otel-instrumentation) or Beeline-specific migration (use beeline-migration).\n","included_files":[{"relative_path":"references/bridge-libraries.md","size_in_bytes":8300},{"relative_path":"references/context-propagation-patterns.md","size_in_bytes":14173},{"relative_path":"references/framework-middleware.md","size_in_bytes":6226},{"relative_path":"references/migration-pitfalls.md","size_in_bytes":6744},{"relative_path":"references/verification-checklist.md","size_in_bytes":5731}],"skill_md_contents":"---\nname: otel-migration\ndescription: >\n  Guide for retrofitting OpenTelemetry into an existing, uninstrumented application.\n  Trigger phrases: \"migrate existing app to OTel\",\n  \"add OpenTelemetry to existing project\", \"retrofit OTel into my codebase\",\n  \"thread context through my code\", \"context propagation\",\n  \"bridge Prometheus metrics to OTel\", \"logging bridge\",\n  \"migrate logging to OTel\", \"slog bridge\", \"logback bridge\",\n  \"verify my instrumentation\", \"traces are disconnected\",\n  \"orphaned spans\", \"migrate to OpenTelemetry\", \"OTel migration plan\",\n  \"how do I sequence an OTel migration\", \"add tracing to existing code\",\n  \"refactor for context propagation\", \"Fiber context gotcha\",\n  \"keep existing logging working with OTel\", \"add OTel without breaking Prometheus\",\n  \"bridge existing metrics\", \"coexist with existing monitoring\",\n  or any request about retrofitting OpenTelemetry into an existing application.\n  This skill is for migrating existing codebases, NOT greenfield instrumentation (use otel-instrumentation)\n  or Beeline-specific migration (use beeline-migration).\nmetadata:\n  version: \"1.0.0\"\n---\n\n# OpenTelemetry Migration for Existing Applications\n\nGuide for retrofitting OpenTelemetry into an existing, uninstrumented application. This covers\nthe phased migration approach, context propagation refactoring, logging and metrics bridges, and\nverification. This is distinct from greenfield OTel setup (see otel-instrumentation skill) and\nBeeline-specific migration (see beeline-migration skill).\n\n## When to Use This Skill\n\nUse this skill when the user has an **existing application** that:\n- Has no OpenTelemetry instrumentation and needs to add it\n- Has existing logging, metrics, or context patterns that must coexist with OTel\n- Needs to refactor function signatures to thread trace context through the call stack\n- Uses a framework with OTel middleware gotchas (e.g., Fiber, Gin, Express)\n\nFor greenfield OTel setup, use the `otel-instrumentation` skill instead.\nFor Beeline-to-OTel migration, use the `beeline-migration` skill instead.\nFor understanding *why* to instrument, see the `observability-fundamentals` skill.\n\n## Migration Phases\n\nThe migration follows six phases in order. Each phase is independently deployable and verifiable.\nContext propagation (Phase 3) is typically ~60% of the effort.\n\n### Phase 1: SDK Initialization and Shutdown\n\nSet up TracerProvider, MeterProvider, and LoggerProvider with OTLP exporters. Wire initialization\nearly in the application's entry point and shutdown in signal handlers.\n\n**Key guidance:**\n- SDK init must happen *before* any application code that might create spans — one of the first\n  things in your entry point, before config loading or storage initialization\n- Shutdown ordering matters: flush traces, then metrics, then logs. Use a timeout (10-30s)\n- If init fails, the application should still work — log the error and continue without telemetry\n- For language-specific SDK setup, consult\n  `${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`\n\n### Phase 2: HTTP Middleware (Auto-Instrumentation)\n\nAdd OTel middleware to your HTTP framework. This gives you automatic spans for every inbound\nrequest with zero code changes to handlers. This is the highest-ROI step.\n\n**Critical:** Different frameworks expose the OTel-enriched context differently. This is the #1\nsource of silent trace breaks. Consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/framework-middleware.md` for\nframework-specific details.\n\n| Framework | How to get OTel context | Common mistake |\n|-----------|------------------------|----------------|\n| Go net/http | `r.Context()` | N/A (standard) |\n| Go Fiber v2 | `c.UserContext()` | Using `c.Context()` (returns fasthttp context without OTel span) |\n| Go Gin | `c.Request.Context()` | Using `c` directly |\n| Go Echo | `c.Request().Context()` | N/A |\n| Python Flask | Automatic (thread-local) | N/A with instrumentation library |\n| Python Django | Automatic (thread-local) | N/A with instrumentation library |\n| Node.js Express | Automatic (AsyncLocalStorage) | N/A with instrumentation library |\n| Java Spring | Automatic (thread-local) | Thread pool context loss |\n| .NET ASP.NET Core | Automatic (AsyncLocal) | N/A |\n| Ruby Rails | Automatic (thread-local) | N/A with instrumentation library |\n\n### Phase 3: Context Propagation Refactoring\n\nThread trace context through your call chain from HTTP handlers (or entry points) down to I/O\noperations. **This is the hardest phase** — typically ~60% of migration effort.\n\nThe difficulty of this phase varies dramatically by language:\n- **Go**: Hardest. Requires adding `context.Context` parameter to every function in the call chain.\n- **Java**: Moderate. Thread-local context propagates automatically within a thread, but breaks\n  across thread pools, CompletableFuture, and reactive streams.\n- **Python**: Easier. `contextvars` propagates automatically within a thread. Pain points are\n  thread pools and multiprocessing.\n- **Node.js**: Easier. `AsyncLocalStorage` propagates through async/await automatically.\n  Pain points are old callback-based code.\n- **.NET**: Easiest. `Activity` propagates through async/await via `AsyncLocal<T>` automatically.\n- **Ruby**: Easier. Thread-local context propagates automatically. Pain with manual thread creation.\n\nFor language-specific patterns and code examples, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/context-propagation-patterns.md`.\n\n### Phase 4: Custom Spans\n\nAdd spans to business logic operations that auto-instrumentation doesn't cover. Defer to the\n`otel-instrumentation` skill for span creation mechanics. Migration-specific guidance:\n\n1. **Start with I/O boundaries** — database calls, external HTTP calls, cache operations\n2. **Then add business logic** — operations that explain *why* time is spent\n3. **Add attributes liberally** — every piece of context makes BubbleUp useful during investigations\n4. **Record outcomes on spans** — result status, error count, duration as attributes\n\nFor attribute naming and span creation patterns, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.\n\n### Phase 5: Logging Migration\n\nReplace or bridge your existing logging library into OTel so logs correlate with traces.\n\n**Key guidance:**\n- You almost certainly want logs going to both stderr (local debugging) AND OTel (trace correlation).\n  This requires a multi-handler/fan-out pattern.\n- OTel log bridges work with structured logging (key-value pairs). If your existing logging uses\n  printf-style format strings, convert to structured format first.\n- Converting from printf-style to structured logging is tedious but mechanical — a good candidate\n  for automated refactoring.\n\nFor language-specific logging bridges and the multi-handler pattern, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/bridge-libraries.md`.\n\n### Phase 6: Metrics Bridge\n\nIf you already have Prometheus metrics (or another metrics library), bridge them to OTel rather\nthan rewriting.\n\n**Key guidance:**\n- Prometheus bridge reads from the existing registry and produces OTel metrics — existing\n  `prometheus.NewCounterVec(...)` calls continue unchanged\n- Keep the Prometheus `/metrics` endpoint if you have existing scrapers. The bridge adds OTLP\n  export *in addition to* scraping.\n- If you want to eventually remove the Prometheus dependency, plan a separate migration later.\n  The bridge buys you time.\n\nFor language-specific metrics bridges, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/bridge-libraries.md`.\n\n## Verification\n\nAfter each phase, verify that instrumentation is correct and complete. Consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/verification-checklist.md` for the\nfull checklist and query patterns.\n\nTo verify locally without a Honeycomb account, use the bundled collector script to capture\nspans as debug output and NDJSON. Consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/local-collector-debug-test.md` for usage, and\n`${CLAUDE_PLUGIN_ROOT}/scripts/start-collector.sh` for the full script.\n\nFor Honeycomb-specific verification queries, also consult the `query-patterns` skill.\n\n## Common Pitfalls\n\nFor a catalog of common mistakes and how to avoid them, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/migration-pitfalls.md`.\n\n## Real-World Calibration\n\nFor reference, a real migration of Gatus (~30k LOC Go, Fiber v2, SQLite/Postgres, Prometheus):\n- **Files changed:** ~45, **Lines:** +840/-645\n- **Effort breakdown:** Context propagation ~60%, Custom spans ~15%, Logging migration ~15%, Everything else ~10%\n- **Bugs encountered:** Fiber `c.Context()` vs `c.UserContext()`, printf-style slog format strings,\n  missing `span.End()` calls, goroutine context reuse\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}