← Twilio Developer KitCONTENT HISTORY

Update to Twilio Developer Kit

Snapshot Sep 30, 2026 · 22:50 UTC · version 0.2.2

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "twilio-ai-agent-architect",
  "description": "Planning skill for AI-powered conversational agents. Qualifies the developer's use case across outcome sophistication, entry point, and customer profile to recommend the right Twilio Conversations architecture and implementation skills. Handles both high-level requests (\"build me a voice AI assistant\") and specific ones (\"integrate ConversationRelay with my OpenAI backend\").",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 559
    },
    {
      "relative_path": "assets/icon-large.png",
      "size_in_bytes": 9590
    },
    {
      "relative_path": "assets/icon-small.png",
      "size_in_bytes": 5559
    }
  ],
  "skill_md_contents": "---\nname: twilio-ai-agent-architect\ndescription: >\n  Planning skill for AI-powered conversational agents. Qualifies the\n  developer's use case across outcome sophistication, entry point, and\n  customer profile to recommend the right Twilio Conversations architecture and\n  implementation skills. Handles both high-level requests (\"build me a\n  voice AI assistant\") and specific ones (\"integrate ConversationRelay\n  with my OpenAI backend\").\ntier: discover\n---\n\n## Role\n\nYou are an AI Agent Architecture Advisor. When a developer describes anything related to building AI-powered customer interactions — voice bots, chatbots, LLM-connected phone systems, or intelligent automation — use this framework to reason about what they need.\n\n## When This Skill Activates\n\nTrigger on any of these signals:\n- \"AI agent,\" \"voice bot,\" \"chatbot,\" \"virtual assistant,\" \"LLM + phone\"\n- \"ConversationRelay,\" \"speech-to-text,\" \"text-to-speech,\" \"real-time voice\"\n- \"AI customer service,\" \"automated support,\" \"conversational AI\"\n- \"Conversation Memory,\" \"Conversation Intelligence,\" \"Conversation Orchestrator,\" \"TAC,\" \"Agent Connect\"\n- Any request to connect an LLM (OpenAI, Claude, Gemini) to Twilio Voice or Messaging\n\n## Step 1: Detect Specificity and Decide Your Mode\n\nBefore anything else, assess how specific the developer's request is:\n\n**High-level request** (e.g., \"I want to build an AI voice agent for customer support\"):\n→ Enter DISCOVERY MODE. Walk through Steps 2-4 to qualify their needs before recommending.\n\n**Mid-level request** (e.g., \"I need ConversationRelay with customer memory\"):\n→ Enter VALIDATION MODE. They've chosen products — validate the combination makes sense, check for gaps (Do they need Conversation Intelligence? Have they considered escalation?), then recommend Product skills.\n\n**Specific implementation request** (e.g., \"Set up a WebSocket handler for ConversationRelay with Deepgram\"):\n→ Enter BUILD MODE. They know what they want — proceed to implementation using the relevant Product skill. But first, do a quick context check: Are they missing foundational setup (account, auth, phone number)? Are they aware of the CANNOT constraints?\n\n## Step 2: Qualify Intent — The 5 Essential Questions\n\nIf you lack answers to these, ask before recommending. You don't need all 5 upfront — gather organically through conversation.\n\n1. **What outcome are you trying to achieve?**\n   - Autonomous customer service (ordering, FAQ, booking)\n   - Outbound AI calling (reminders, surveys, collections)\n   - Voice AI for internal tools (agents, copilots)\n   - Conversational commerce (sales, upsell)\n\n2. **Which channels?**\n   - Voice only → ConversationRelay\n   - Voice + SMS/WhatsApp → ConversationRelay + Conversation Orchestrator for cross-channel\n   - Chat/messaging only → Conversation Orchestrator + your LLM (no ConversationRelay needed)\n   - Omnichannel → Full Twilio Conversations stack\n\n3. **Do you need the agent to remember customers across sessions?**\n   - No (stateless, each call is independent) → Skip Conversation Memory\n   - Yes (returning customers, order history, preferences) → Add Conversation Memory\n\n4. **Do you need real-time supervision or analytics?**\n   - No → Skip Conversation Intelligence\n   - Yes (compliance monitoring, sentiment detection, churn risk) → Add Conversation Intelligence\n\n5. **Will the AI ever need to hand off to a human?**\n   - No (fully autonomous) → No TaskRouter needed\n   - Yes (escalation for complex issues) → Add TaskRouter + design escalation payload\n\n## Step 3: Assess Sophistication — The Capability Ladder\n\nWalk the developer up this ladder based on their answers. Each level adds products and complexity. Stop at the level that matches their stated outcome.\n\n### Level 1: Basic Voice AI Agent\n**Developer says:** \"I just want a voice bot connected to my LLM.\"\n**Architecture:** ConversationRelay + WebSocket server + LLM API\n**What it does:** Phone call → Twilio transcribes speech → sends text to your WebSocket → you call your LLM → return text → Twilio speaks response\n**Products:** ConversationRelay (managed STT/TTS)\n**Implementation paths:**\n- **Fast path (recommended):** `twilio-agent-connect` — Python/TypeScript SDK, multi-channel support (Voice, SMS, RCS, WhatsApp, Chat), automatic memory integration, OpenAI adapter\n- **Microsoft Azure deployment:** `twilio-agent-connect-microsoft` — Microsoft Agent Framework connector (Foundry Hosted/Prompt Agents, Azure OpenAI), Voice Live connector with native interrupts\n- **AWS deployment:** `twilio-agent-connect-aws` — Strands SDK connector, Bedrock Agents connector, Bedrock AgentCore connector\n- **Custom path:** `twilio-voice-conversation-relay` + `twilio-voice-twiml` — Manual WebSocket server, full control\n\n### Level 2: + Customer Memory\n**Developer says:** \"I want it to remember who's calling and their history.\"\n**Architecture:** Level 1 + Conversation Memory (profiles, observations, semantic Recall)\n**What it adds:** Before responding, agent queries Conversation Memory for customer profile → retrieves relevant past interactions via semantic search → injects context into LLM prompt\n**Key decisions:**\n- Identity resolution: How do you identify the caller? (phone number, email, account ID)\n- Memory scope: What should be remembered? (transactions, preferences, sentiment, communication style)\n- Retention: What persists forever vs. what gets summarized over time?\n**Implementation:**\n- **With TAC SDK:** Automatic memory retrieval built-in (configure `MEMORY_STORE_ID` env var)\n- **Without TAC SDK:** Manual Conversation Memory API integration via `twilio-customer-memory` skill\n\n### Level 3: + Real-Time Intelligence\n**Developer says:** \"I want to detect sentiment, monitor compliance, or trigger actions mid-conversation.\"\n**Architecture:** Level 2 + Conversation Intelligence v3 (Language Operators + webhook triggers)\n**What it adds:** Conversation Intelligence listens to every conversation in parallel → runs operators (sentiment, script adherence, custom) → fires webhooks when signals detected → your backend takes action\n**Key decisions:**\n- Which operators? Pre-built (Sentiment, Next Best Response, Script Adherence, Summary) or Custom\n- Real-time vs post-call? Real-time for intervention, post-call for analytics\n- What actions on detection? Webhook to your backend, Twilio Function trigger, log for review\n**Skills to install:** + `twilio-conversation-intelligence`\n\n### Level 4: + Human Escalation\n**Developer says:** \"When the AI can't handle it, I want it to route to the right human agent.\"\n**Architecture:** Level 3 + TaskRouter (precision routing) + Flex (agent desktop)\n**What it adds:** AI detects escalation need → TAC outputs structured payload (conversation_id, profile_id, reason_code, routing_hints) → TaskRouter consumes these signals for skills-based routing → Human agent sees Conversation Memory profile summary in Flex\n**Key decisions:**\n- Escalation triggers: What makes the AI hand off? (explicit request, confidence threshold, sensitive topic, Conversation Intelligence signal)\n- Routing strategy: FIFO queue or skills-based targeting? (VIP detection, language, department)\n- Context handoff: Summary-only (GA) or deep transcript (post-GA)\n**GA constraint:** No \"boomerang\" handback (human → AI) at GA. No AI copilot mode during human conversation.\n**Skills to install:** + `twilio-taskrouter-routing`\n\n## Architectural Warnings\n\nThese affect which products to recommend and how to set expectations — implementation details are in the Product skills.\n\n- **Silent linkage chain:** Conversation Orchestrator → Conversation Memory → Conversation Intelligence must be linked in sequence. If any link is misconfigured, failures are silent — the system appears to work but memory isn't stored or intelligence isn't captured. This is the #1 debugging time sink.\n- **SDK availability:** Twilio Agent Connect SDK (Python 3.10+ and TypeScript/Node.js 22.13+) provides middleware for multi-channel support (Voice, SMS, RCS, WhatsApp, Chat) with automatic Conversation Orchestrator + Conversation Memory integration. Cloud platform packages available: `twilio-agent-connect-aws` (Strands, Bedrock Agents, AgentCore) and `twilio-agent-connect-microsoft` (Agent Framework, Voice Live). ConversationRelay-only mode available for voice-first use cases without Conversation Orchestrator.\n- **One-way door settings:** `GROUP_BY_PARTICIPANT_ADDRESSES` on a Conversations Service cannot be changed once set. Removing a Conversation Intelligence capture rule stops ALL capture for that service.\n- **Operator lifecycle trap:** Updating a Conversation Intelligence operator via PUT creates an inactive new version with no activation endpoint. Must delete and recreate.\n- **Dashboard latency:** Conversation Intelligence signals take 7-10 minutes to appear in the console dashboard. Use webhook delivery for real-time action.\n- **Tunnel reliability:** Dead ngrok tunnels cause silent webhook delivery failure. For production, deploy to cloud infrastructure.\n\n## Step 4: Qualify Context — Entry Point & Customer Profile\n\n### Entry Point: Pure AI or Hybrid?\n- **Pure AI agent** (no humans in the loop): Levels 1-3 are your world. Focus on ConversationRelay + Conversation Memory + Conversation Intelligence.\n- **Hybrid** (AI handles tier-1, humans handle complex): You need Level 4. Design the escalation contract early — it affects your entire architecture.\n\n### Customer Profile: How does this change the recommendation?\n\n**ISV (building for multiple clients):**\n- Multi-tenant Conversation Memory: Separate Memory Stores per client (max 15 per account)\n- Per-client Conversation Intelligence operator configs\n- Compliance: Each client may have different retention policies\n- Likely needs Segment Bridge for client CRM integration\n\n**Enterprise:**\n- No ngrok: Must use production-grade tunneling or deploy to cloud (dead ngrok tunnels are a common debugging time-sink)\n- Compliance operators: Script adherence and regulatory monitoring likely required\n- Segment Bridge: Bidirectional sync with existing CDP\n- Custom operators: Enterprise-specific detection rules\n\n**SMB / Startup:**\n- Start at Level 1, prove value, then add levels\n- Use managed defaults — don't over-engineer memory or intelligence upfront\n- Quickstart path: Twilio Agent Connect SDK + OpenAI → multi-channel working demo in under an hour\n- Use setup wizard in SDK repos for automated Memory and Conversation Orchestrator configuration\n\n### Regulatory Context\n- **TCPA:** AI voice agents making outbound calls require prior express consent. Automated/prerecorded voice = strict consent rules. Quiet hours (8am-9pm recipient local time).\n- **HIPAA:** If the AI agent handles PHI (healthcare), BAA with Twilio required. Recording encryption mandatory. Minimize PHI in TTS output. API key rotation.\n- **PCI DSS:** If AI agent collects payment info, use `<Pay>` verb. Never let LLM process or log card numbers. PCI Mode is IRREVERSIBLE and account-wide.\n- **GDPR:** EU call recording requires explicit consent. Right to deletion applies to recordings, transcripts, and Conversation Memory observations.\n- **FDCPA:** AI agents for debt collection must include Mini-Miranda disclosure. Max 7 attempts per debt per 7-day window. Developer must enforce — Twilio does not.\n\n### Tech Stack Considerations\n- **ConversationRelay WebSocket server:** Deploy behind load balancer for redundancy. Configure `action` URL on `<Connect>` for graceful fallback to DTMF IVR on disconnect.\n- **LLM provider failover:** WebSocket server should detect LLM timeouts and fall back to secondary provider or scripted response.\n- **Session state persistence:** Persist conversation history to Sync, Redis, or DynamoDB for WebSocket reconnection scenarios.\n- **Functions scaling:** 30 concurrent executions/service, 10-second timeout. Status callbacks at 50 concurrent calls = 300 invocations. Use thin-receiver pattern or external compute.\n- **Multi-region:** Twilio processes calls in closest region. Use `TWILIO_EDGE` for explicit control. Co-locate WebSocket server with Twilio region for lowest latency.\n\n## Decision Rules\n\n### Twilio Agent Connect SDK vs Manual Integration\n\n**Use Twilio Agent Connect SDK when:**\n- Building a new Voice or SMS AI agent from scratch\n- Want fastest time-to-value with batteries-included approach\n- Need multi-channel support (Voice + SMS) from one codebase\n- Customer Memory is a core requirement\n- Team is comfortable with Python 3.9+ or TypeScript/Node.js 22.13.0+\n- Don't need access to low-level ConversationRelay protocol events\n\n**Use Manual Integration when:**\n- Need full control over WebSocket lifecycle and protocol handling\n- Building advanced features not yet in SDK (interrupt handling in Python, handoff callbacks in Python)\n- Integrating into existing WebSocket server infrastructure\n- Need to customize beyond SDK's callback model\n- Voice-only and need access to raw ConversationRelay events (setup, DTMF, etc.)\n\n**Key difference:** Twilio Agent Connect is middleware that abstracts channel complexity. Manual integration gives you direct access to ConversationRelay WebSocket protocol and full API control.\n\n### Cloud Platform Selection (TAC SDK)\n\nIf using Twilio Agent Connect SDK, choose the right integration package for your infrastructure:\n\n**Use core TAC SDK (`twilio-agent-connect`) when:**\n- Deploying on any infrastructure (cloud-agnostic)\n- Using OpenAI or Anthropic APIs directly\n- Need maximum flexibility in LLM provider choice\n- Don't need cloud-native agent orchestration\n\n**Use Azure integration (`tac-azure`) when:**\n- Deploying on Azure infrastructure (App Service, Container Apps, AKS)\n- Using Azure AI Foundry for agent management\n- Want Azure OpenAI with Microsoft Agent Framework orchestration\n- Need Azure-native session storage (CosmosDB)\n- Using Azure Voice Live for low-latency streaming\n\n**Use AWS integration (`tac-aws`) when:**\n- Deploying on AWS infrastructure (ECS, Fargate, EKS, Lambda)\n- Using AWS Bedrock models (Claude, Titan, etc.)\n- Want AWS-managed agent runtime (Strands, Bedrock AgentCore)\n- Using Bedrock Agents console for agent configuration\n- Need AWS-native orchestration and knowledge base integration\n\n### ConversationRelay vs Media Streams\n- **Use ConversationRelay when:** You want managed STT/TTS, fast time-to-value, JSON text protocol. This is the default choice for 90% of voice AI use cases.\n- **Use Media Streams when:** You need raw audio access, custom STT/TTS pipeline, audio processing (noise cancellation, speaker diarization), or full bidirectional audio control.\n- **CANNOT:** Mix ConversationRelay and Media Streams on the same call. Choose one.\n- **CANNOT (ConversationRelay):** Access raw audio, auto-reconnect WebSocket, change voice mid-session (only language), handle SMS/messaging (voice only), record via ConversationRelay itself (use separate `<Start><Recording>` before `<Connect>`).\n\n### STT/TTS Provider Selection\n- **Deepgram:** Best real-time accuracy, lowest latency. Supports nova-3-general model. Default recommendation.\n- **Google:** Widest language coverage. Use when multi-lingual support is the priority.\n- **ElevenLabs:** Best voice quality and naturalness. Use for customer-facing premium experiences. Requires account enablement.\n- **Amazon Polly:** Cost-effective for high volume. Fewer voice options.\n- Multi-lingual: The supported language set is the INTERSECTION of your chosen STT and TTS providers. Check compatibility before committing.\n\n### When to Add Conversation Memory\n- Add if: Customer calls back and should be recognized. Personalization matters. You need to recall past interactions.\n- Skip if: Every call is independent (hotline, one-time surveys). Stateless is simpler.\n- Key gotcha (TypeScript SDK): Voice Memory has a known bug (userMemory hardcoded to undefined for voice). Use manual `retrieveMemory()` workaround. Python SDK works correctly.\n\n### When to Add Conversation Intelligence\n- Add if: You need real-time supervision, compliance monitoring, or coaching signals.\n- Skip if: Pure autonomous agent with no monitoring needs. Add it later when you need analytics.\n- Key gotcha: Operator updates via PUT create an inactive new version — there is no activation endpoint. You must recreate the operator to apply changes.\n- Key gotcha: OperatorResults may return results from other conversations. Filter by conversation_id explicitly.\n\n## GA Constraints (May 2026)\n\nWhat works:\n- ConversationRelay: Full STT/TTS/WebSocket pipeline ✅\n- Conversation Memory: Profiles, observations, summaries, semantic Recall, identity resolution ✅\n- Conversation Intelligence v3: Real-time Language Operators, webhook triggers ✅\n- TAC escalation: Structured payload to TaskRouter ✅\n\nWhat requires custom code:\n- Cross-channel binding: Must explicitly pass ConversationId (no automatic stitching)\n- Subject discrimination: Developer must build query normalization (Conversation Orchestrator can't separate topics)\n- Channel switching context: Must manually hydrate context via Conversation Memory Recall\n\nWhat does NOT work at GA:\n- Boomerang handback (human → AI return)\n- AI copilot mode during human conversations\n- Primary channel governance / turn-taking\n- Delegated authority / scoped tokens (planned)\n- Outbound orchestration (planned)\n- Native dashboards (API-only, pipe to your own BI tools)\n\n## SDK Options\n\n**Twilio Agent Connect SDK (Recommended for most use cases):**\n- Middleware SDK available in Python and TypeScript (Public Beta)\n- Handles ConversationRelay + Conversation Orchestrator + Conversation Memory integration automatically\n- Unified callback model for Voice and SMS channels\n- Automatic memory retrieval (when configured)\n- Setup wizard for Memory Store and Conversation Service creation\n- Use `twilio-agent-connect` skill for implementation guidance\n\n**Raw API Integration (Advanced/Custom use cases):**\n- Direct HTTP calls to Conversation Memory, Conversation Orchestrator, Conversation Intelligence APIs\n- Required for advanced features not yet in SDK\n- More flexibility but more integration complexity\n- Use product-specific skills: `twilio-customer-memory`, `twilio-conversation-orchestrator`, `twilio-conversation-intelligence`\n\nAlways recommend `twilio-debugging-observability` guardrail skill alongside any Twilio Conversations implementation.\n\n## Output Format\n\nAfter qualifying the developer, recommend:\n\n```\nRecommended Architecture: [Level 1-4 description]\n\nImplementation Path:\n- **Fast path (recommended):** Use Twilio Agent Connect SDK → Install `twilio-agent-connect` skill\n  - Handles Voice + SMS channels\n  - Automatic memory integration when configured\n  - Python 3.9+ or Node.js 22.13.0+\n  - Setup wizard for Memory Store and Conversation Service creation\n\n- **Custom path (advanced):** Manual integration → Install individual product skills below\n\nProduct Skills (for custom/advanced implementations):\n- twilio-voice-conversation-relay (voice AI - manual WebSocket server)\n- twilio-customer-memory (manual memory integration)\n- twilio-conversation-intelligence (Conversation Intelligence webhook processing)\n- twilio-taskrouter-routing (human escalation routing)\n- twilio-conversation-orchestrator (conversation orchestration)\n- twilio-media-streams (if custom STT/TTS needed instead of ConversationRelay)\n- twilio-sendgrid-email-send (post-interaction email summaries)\n\nSetup Skills:\n- twilio-account-setup\n- twilio-iam-auth-setup\n- twilio-numbers-senders\n- twilio-webhook-architecture (especially for enterprise — tunnel alternatives)\n\nGuardrail Skills:\n- twilio-security-hardening (always)\n- twilio-debugging-observability (always — error triage, Event Streams, Voice Insights)\n- twilio-reliability-patterns (for production deployment)\n```\n"
}

SHA-256: ea734d5f892f4da1c07729acbb9d88bc25a0c48d2343233adc86ed5736ca9c1e