← Plugin catalog
Productivity
ElevenLabs
Eleven Labs Inc. v1.0.0
Publisher description
From the marketplace listing
ElevenLabs helps users create and manage voice agents, review conversations, manage knowledge and tools, and generate speech, images, video and custom voices.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package83 files · 201 KBBrowse files →
Skill instructions
agents24.3 KB
---
name: agents
description: Build voice AI agents with ElevenLabs. Use when creating voice assistants, customer service bots, interactive voice characters, or any real-time voice conversation experience, and when configuring an agent's tools, workflows, or procedures, including creating, editing, compiling, and publishing procedure drafts on an agent branch over the SDKs or REST API.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Agents Platform
Build voice AI agents with natural conversations, multiple LLM providers, custom tools, and easy web embedding.
> **Setup:** See [Installation Guide](references/installation.md) for CLI and SDK setup.
## Quick Start with CLI
The ElevenLabs CLI is the recommended way to create and manage agents:
```bash
# Install CLI and authenticate
npm install -g @elevenlabs/cli
elevenlabs auth login
# Initialize project and create an agent
elevenlabs agents init
elevenlabs agents add "My Assistant" --template complete
# Push to ElevenLabs platform
elevenlabs agents push
```
**Available templates:** `complete`, `minimal`, `voice-only`, `text-only`, `customer-service`, `assistant`
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
agent = client.conversational_ai.agents.create(
name="My Assistant",
conversation_config={
"agent": {
"first_message": "Hello! How can I help?",
"language": "en",
"prompt": {
"prompt": "You are a helpful assistant. Be concise and friendly.",
"llm": "gemini-2.0-flash",
"temperature": 0.7
}
},
"tts": {"voice_id": "JBFqnCBsd6RMkjVDRZzb"}
}
)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const client = new ElevenLabsClient();
const agent = await client.conversationalAi.agents.create({
name: "My Assistant",
conversationConfig: {
agent: {
firstMessage: "Hello! How can I help?",
language: "en",
prompt: {
prompt: "You are a helpful assistant.",
llm: "gemini-2.0-flash",
temperature: 0.7
}
},
tts: { voiceId: "JBFqnCBsd6RMkjVDRZzb" }
}
});
```
### CLI
The CLI reads `ELEVENLABS_API_KEY` from the environment automatically:
```bash
elevenlabs agents create \
--json '{"name": "My Assistant", "conversation_config": {"agent": {"first_message": "Hello!", "language": "en", "prompt": {"prompt": "You are helpful.", "llm": "gemini-2.0-flash"}}, "tts": {"voice_id": "JBFqnCBsd6RMkjVDRZzb"}}}'
```
## Starting Conversations
### Temporary LiveKit WebSocket Pin
Until the ElevenLabs LiveKit server supports `/rtc/v1`, browser clients using WebRTC can fail or stall during the underlying LiveKit WebSocket handshake with `livekit-client` versions newer than `2.16.1`. For React, Next.js, Electron, or other `@elevenlabs/client` / `@elevenlabs/react` integrations that use `connectionType: "webrtc"` or hit `wss://livekit.rtc.elevenlabs.io/rtc/v1`, add this temporary pin to `package.json`:
```json
{
"overrides": {
"livekit-client": "2.16.1"
}
}
```
Use the pin when the app logs `/rtc/v1` 404s, `v1 RTC path not found`, or `could not establish pc connection` during session startup. This is a LiveKit server compatibility workaround for WebRTC sessions, not the ElevenLabs `connectionType: "websocket"` transport. Remove it after the upstream LiveKit server or SDK issue is fixed.
**Authenticated WebRTC:** Request a session token from your backend. The response includes both
the token and the conversation ID:
```python
session = client.conversational_ai.conversations.get_webrtc_token(
agent_id="your-agent-id",
)
print(session.token, session.conversation_id)
```
**Server-side (Python):** Get signed URL for client connection:
```python
signed_url = client.conversational_ai.conversations.get_signed_url(
agent_id="your-agent-id",
environment="staging",
)
```
**Client-side (JavaScript):**
```javascript
import { Conversation } from "@elevenlabs/client";
const conversation = await Conversation.startSession({
agentId: "your-agent-id",
environment: "staging",
overrides: { asr: { keywords: ["ElevenLabs", "TechCorp"] } },
onMessage: (msg) => console.log("Agent:", msg.message),
onUserTranscript: (t) => console.log("User:", t.message),
onPing: (event) => console.log("Estimated latency:", event.ping_ms),
onError: (e) => console.error(e)
});
```
**React Hook:** Wrap hook consumers in `ConversationProvider`. Prefer granular hooks such as
`useConversationControls` and `useConversationStatus` for session controls and UI state;
`useConversation` remains available as the convenience all-in-one hook. Pass provider-level
callbacks such as `onError` when you want React to handle conversation errors in one place.
```typescript
import {
ConversationProvider,
useConversationControls,
useConversationStatus,
} from "@elevenlabs/react";
function Agent({ signedUrl }: { signedUrl: string }) {
const { startSession, endSession } = useConversationControls();
const { status } = useConversationStatus();
if (status === "connected") {
return <button onClick={endSession}>End conversation</button>;
}
return (
<button onClick={() => startSession({ signedUrl })}>
Start conversation
</button>
);
}
function App({ signedUrl }: { signedUrl: string }) {
return (
<ConversationProvider
onError={(error) => console.error("Conversation error:", error)}
onPing={(event) => console.log("Estimated latency:", event.ping_ms)}
>
<Agent signedUrl={signedUrl} />
</ConversationProvider>
);
}
```
## Configuration
| Provider | Models |
|----------|--------|
| OpenAI | `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5`, `gpt-5.5-2026-04-23`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5.4-2026-03-05`, `gpt-5.4-mini-2026-03-17`, `gpt-5.4-nano-2026-03-17`, `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `gpt-4.1`, `gpt-4.1-mini`, `gpt-4.1-nano`, `gpt-4o`, `gpt-4o-mini`, `gpt-4-turbo` |
| Anthropic | `claude-opus-4-7`, `claude-sonnet-4-6`, `claude-sonnet-4-5`, `claude-sonnet-4`, `claude-haiku-4-5`, `claude-3-7-sonnet`, `claude-3-5-sonnet`, `claude-3-haiku` |
| Google | `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.1-flash-lite-preview`, `gemini-3.1-pro-preview`, `gemini-3-pro-preview`, `gemini-3-flash-preview`, `gemini-2.5-flash`, `gemini-2.5-flash-lite`, `gemini-2.0-flash`, `gemini-2.0-flash-lite` |
| ElevenLabs | `glm-45-air-fp8`, `qwen3-30b-a3b`, `qwen36-35b-a3b`, `qwen35-35b-a3b`, `qwen35-397b-a17b`, `gpt-oss-120b` |
| Custom | `custom-llm` (bring your own endpoint) |
Use `GET /v1/convai/llm/list` to inspect the current model catalog, including deprecation state, token/context limits, capability flags such as image-input support, and model-specific reasoning effort support.
**Popular voices:** `JBFqnCBsd6RMkjVDRZzb` (George), `EXAVITQu4vr4xnSDxMaL` (Sarah), `onwK4e9ZLuTAKqWW03F9` (Daniel), `XB0fDUnXU5powFXDhCwa` (Charlotte)
**Turn eagerness:** `patient` (waits longer for user to finish), `normal`, or `eager` (responds quickly)
See [Agent Configuration](references/agent-configuration.md) for all options.
## System Prompt Structure
Section the prompt with markdown headings — the model prioritizes and interprets instructions more reliably ([prompting guide](https://elevenlabs.io/docs/eleven-agents/best-practices/prompting-guide)):
```
# Personality – named character, 2-3 traits
# Environment – where they work, who they talk to
# Tone – vocal style as 4-5 bullets
# Goal – what success looks like (numbered for multi-step flows)
```
Keep instructions short and action-based. Mark critical steps with "This step is important." For critical refusal/safety rules, include concise instructions in the prompt and also configure independent custom Guardrails via `platform_settings.guardrails` (see [Guardrails](#guardrails)).
## Tools
Extend agents with webhook, client, or built-in system tools. Tools are defined inside `conversation_config.agent.prompt`:
Workspace environment variables can resolve per-environment server tool URLs, headers, and auth connections, and runtime system variables such as `{{system__conversation_history}}` can pass full conversation context into tool calls when needed.
```python
"prompt": {
"prompt": "You are a helpful assistant that can check the weather.",
"llm": "gemini-2.0-flash",
"tools": [
# Webhook: server-side API call
{"type": "webhook", "name": "get_weather", "description": "Get weather",
"api_schema": {"url": "https://api.example.com/weather", "method": "POST",
"request_body_schema": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}}},
# Client: runs in the browser
{"type": "client", "name": "show_product", "description": "Display a product",
"parameters": {"type": "object", "properties": {"productId": {"type": "string"}}, "required": ["productId"]}}
],
"built_in_tools": {
"end_call": {},
"transfer_to_number": {"transfers": [{"transfer_destination": {"type": "phone", "phone_number": "+1234567890"}, "condition": "User asks for human support"}]},
"start_procedure": {}
}
}
```
**Client tools** run in browser:
```javascript
clientTools: {
show_product: async ({ productId }) => {
document.getElementById("product").src = `/products/${productId}`;
return { success: true };
}
}
```
See [Client Tools Reference](references/client-tools.md) for complete documentation.
### Built-in System Tools
Set under `conversation_config.agent.prompt.built_in_tools`. `{}` enables defaults; provide `description` to customize; omit to disable.
| Tool | Enable for |
|------|------------|
| `end_call` | All agents |
| `language_detection` | Multilingual agents |
| `transfer_to_number` | Phone-based human escalation |
| `transfer_to_agent` | Multi-agent workflows |
| `start_procedure` | Procedure-guided conversations (see [Procedures](#procedures)) |
| `end_procedure` | Completing active procedures |
| `skip_turn` | Tutoring / coaching (silent listening) |
| `voicemail_detection` | Outbound calling |
| `play_keypad_touch_tone` | IVR navigation |
`run_subagent` is a system tool for delegating a task to another configured agent. Add it to
`conversation_config.agent.prompt.tools` with `params.system_tool_type: "run_subagent"` and an
`agents` array. Each entry requires `agent_id` and `description`; `branch_id` and a JSON-schema
`parameters` object are optional.
`knowledge_base` is a system tool for letting the model choose how to inspect attached knowledge.
Add it to `conversation_config.agent.prompt.tools` with `type: "system"`, a `name`, and
`params.system_tool_type: "knowledge_base"`. Use `enabled_strategies` to expose any combination of
`cat`, `keyword`, `semantic`, and `ls`:
```json
{
"type": "system",
"name": "knowledge_base",
"description": "Search the attached knowledge base.",
"params": {
"system_tool_type": "knowledge_base",
"enabled_strategies": ["semantic", "keyword"]
}
}
```
### Integration Tools
Pre-built connectors managed by the platform. Create a connection with credentials, then attach via `tool_ids`:
| Integration | Use case |
|-------------|----------|
| `calcom` | Scheduling appointments |
| `salesforce` | CRM lookups, case creation |
| `hubspot` | CRM, marketing, contacts |
| `zendesk` | Support ticketing |
Three-step flow: `POST /v1/convai/api-integrations/{id}/connections` → `GET /v1/convai/api-integrations/{id}/tools` → `POST /v1/convai/tools` with `api_integration_id` and `api_integration_connection_id`. Attach to the agent with `"prompt": {"tool_ids": ["tool_xxxx"]}`. Inline `tools` and `tool_ids` can coexist — prefer an integration over a duplicate custom webhook.
### Public-API Webhook Examples
No-auth APIs useful for prototypes (URLs must be HTTPS):
| Tool | URL | Purpose |
|------|-----|---------|
| `get_weather` | `https://wttr.in/{location}?format=j1` | Current weather |
| `search_wikipedia` | `https://en.wikipedia.org/api/rest_v1/page/summary/{topic}` | Topic summary |
| `get_exchange_rate` | `https://open.er-api.com/v6/latest/{base_currency}` | FX rates |
## Workflows
Route conversations through discrete steps with branching logic. Define under the agent's top-level `workflow` field. Reference: [Agent Workflows](https://elevenlabs.io/docs/eleven-agents/customization/agent-workflows).
**Node types:** `start` (ID must be `"start_node"`), `end`, `override_agent` (subagent step with `label` + `additional_prompt`), `dispatch_tool` (executes a tool with success/failure routing), `agent_transfer`, `transfer_to_number`.
**Edge types:** `unconditional`, `llm` (natural-language condition), `expression` (deterministic data check). Tool nodes have separate success/failure edges.
**Scope tools per step** with `additional_tool_ids` on a node — prevents the wrong tool firing at the wrong step. Set `additional_tool_ids: []` on conversational routing nodes such as greeting and `classify_intent` so they only converse:
```json
{
"type": "override_agent",
"label": "Book Appointment",
"additional_prompt": "Discuss preferred dates and doctors. Show the booking form once agreed.",
"entry_behavior": "wait_for_user",
"additional_tool_ids": ["show_booking_form", "display_appointment_card"],
"position": {"x": 0, "y": 400}
}
```
Include `position` (`{x, y}`) on every node so the editor renders cleanly. Start at `y=0`, put `end` at the bottom, and space branches horizontally at `x=-150` and `x=150`; suggested spacing is 200px vertical between levels and 300px horizontal between branches. Keep workflows to 4-7 nodes and always have a path to `end`.
Use `entry_behavior` on `override_agent` nodes to choose whether a sub-agent speaks immediately (`generate_immediately`), waits for user input (`wait_for_user`), or lets the platform decide (`auto`).
For nested agent transfers, set `enable_nesting` on a `standalone_agent` node and
`return_when_nested` on an `end` node that should return control to the parent workflow.
## Procedures
Reusable instruction blocks an agent runs when a trigger matches. A procedure is `free_form` (markdown guidance the agent adapts, and the only type that can reference knowledge base documents) or `deterministic` (ordered, typed steps for flows that must run consistently). Procedures are in Alpha. See [Using the Procedure API](references/using-procedure-api.md) for the full CLI and SDK flow, and [Writing Procedures](references/writing-procedures.md) for the step schema and authoring rules.
Procedures live on an agent branch, and every write stages a per-user draft:
| Operation | Call |
|-----------|------|
| List, create, read, update, discard, remove | `/v1/convai/agents/{agent_id}/branches/{branch_id}/procedures...` (`procedures.*` and `procedures.drafts.*` in the SDKs) |
| Compile | `POST .../procedures/compile` (`procedures.compile`) |
| Publish | `PATCH /v1/convai/agents/{agent_id}?branch_id=...` (`agents.update`) |
Semantics worth knowing before writing any of these calls:
- Nothing reaches the live agent until you publish. Publishing is not a procedure endpoint; one PATCH on the agent versions every changed procedure draft on the branch.
- `GET .../procedures/{procedure_id}` reads branch HEAD and returns `404` until that procedure's first publish. Read the `/draft` variant to see a procedure you just created; do not retry the create.
- Compile only when structured (`deterministic`) procedures changed. Compilation turns them into workflow nodes, so the publish must carry the `workflow` that compile returned. Free-form-only changes publish without compiling, because the agent loads free-form procedures from their published versions.
- Compile validates structured content and is the only way to check it. On `400` it returns `errors` keyed by procedure ID with the offending field `path`; repair the draft and compile again rather than publishing.
- A draft update replaces the whole body. Read the draft first, then resend `name`, `type`, and `trigger` alongside the new `content`.
- `content` is markdown for a `free_form` procedure, and a JSON-encoded object with a `trigger` and a `steps` array for a `deterministic` one. Serialize it; do not hand-escape quotes.
- Routing is driven by the `trigger` text, not the procedure name. Write concrete, non-overlapping triggers that cover the phrasings a user would actually say.
- Procedure APIs require `elevenlabs` (Python) or `@elevenlabs/elevenlabs-js` at `2.60.0` or newer.
## Guardrails
Layered safety enforcement that runs independently of the LLM — configured under `platform_settings.guardrails`, not in the system prompt. Reference: [Guardrails](https://elevenlabs.io/docs/eleven-agents/best-practices/guardrails).
```json
"platform_settings": {
"guardrails": {
"version": "1",
"focus": {"is_enabled": true},
"prompt_injection": {"is_enabled": true},
"content": {"config": {"harassment": {"is_enabled": true, "threshold": 0.5}}},
"custom": {
"config": {
"configs": [{
"is_enabled": true,
"name": "No medical diagnoses",
"prompt": "Block the agent from providing medical diagnoses or treatment advice.",
"execution_mode": "blocking",
"model": "gemini-2.5-flash-lite",
"history_message_count": 1,
"trigger_action": {"type": "retry", "feedback": "Reason: {{trigger_reason}}"}
}]
}
}
}
}
```
**Types:** `focus` (on-topic), `prompt_injection` (manipulation defense), `content` (category filters), `custom` (LLM-evaluated domain rules). Content categories include `harassment`, `profanity`, `sexual`, `violence`, `self_harm`, and `medical_and_legal_information` — threshold range `0.0`–`1.0` (default `0.3`). Custom rules use `execution_mode: "blocking"` with a `model`, `history_message_count`, and `trigger_action` (e.g., `retry` with feedback). Custom guardrails evaluate in parallel and fail-open.
**Per vertical:** healthcare/finance/legal → enable `medical_and_legal_information`; education/youth → `sexual`/`violence`/`self_harm`/`profanity`; support/sales → `harassment`/`profanity`. All agents benefit from `focus` + `prompt_injection` + 2-4 custom rules.
## Testing Agents
Three test types via `POST /v1/convai/agent-testing/create`, then attached with PATCH on the agent. Reference: [Agent Testing](https://elevenlabs.io/docs/eleven-agents/customization/agent-testing).
| Type | Purpose |
|------|---------|
| `llm` | Scenario test — does the agent respond appropriately to a message? |
| `tool` | Tool-call test — right tool, right parameters? |
| `simulation` | Multi-turn flow with a simulated user persona |
```json
// Tool-call test (snake_case throughout; chat_history role is "user" or "agent")
{
"name": "Books with correct doctor and date",
"type": "tool",
"chat_history": [
{"role": "user", "message": "Dr. Smith on March 5 at 2pm", "time_in_call_secs": 10}
],
"tool_call_parameters": {
"referenced_tool": {"id": "show_booking_form", "type": "client"},
"parameters": [
{"path": "doctor_name", "eval": {"type": "llm", "description": "Should reference Dr. Smith"}},
{"path": "date", "eval": {"type": "regex", "pattern": "2025-03-05|March 5"}}
]
}
}
```
Eval strategies: `exact`, `regex`, `llm`. Prompt evaluation criteria can use binary scoring or
numeric scoring with `scoring_mode: "numeric_uniform"`, `max_score`, and `score_instructions`;
numeric scores are normalized into the aggregate conversation success percentage. Attach via an agent update:
```bash
elevenlabs agents update --agent-id "your-agent-id" \
--json '{"platform_settings": {"testing": {"attached_tests": [{"test_id": "test_xxxx"}]}}}'
```
Run selected tests with `POST /v1/convai/agents/{agent_id}/run-tests`. The request
body requires `tests` and accepts `repeat_count` from `1` to `50` for repeated runs.
Simulation tests can define up to 30 `success_conditions` prompts; all criteria are
evaluated and merged into the final result.
Simulation tests can also define `tool_mock_overrides`, keyed by tool ID, to replace shared response
mocks for one test. Each override is an array of mocks with a required `mock_result`; set
`is_error: true` to exercise a tool-failure path. Overrides only apply to tools enabled for mocking
through `tool_mock_config`.
For completed conversations, rerun one evaluation criterion with `POST /v1/convai/conversations/{conversation_id}/analysis/evaluations/run` and a request body containing `evaluation_id`.
## Widget Embedding
```html
<elevenlabs-convai agent-id="your-agent-id"></elevenlabs-convai>
<script src="https://unpkg.com/@elevenlabs/convai-widget-embed" async type="text/javascript"></script>
```
Customize with attributes: `avatar-image-url`, `action-text`, `start-call-text`, `end-call-text`.
See [Widget Embedding Reference](references/widget-embedding.md) for all options.
## Outbound Calls
Make outbound phone calls using your agent via Twilio or Exotel integration:
The examples below use Twilio. See the reference for Exotel usage.
### Python
```python
response = client.conversational_ai.twilio.outbound_call(
agent_id="your-agent-id",
agent_phone_number_id="your-phone-number-id",
to_number="+1234567890",
call_recording_enabled=True
)
print(f"Call initiated: {response.conversation_id}")
```
### JavaScript
```javascript
const response = await client.conversationalAi.twilio.outboundCall({
agentId: "your-agent-id",
agentPhoneNumberId: "your-phone-number-id",
toNumber: "+1234567890",
callRecordingEnabled: true,
});
```
### CLI
```bash
elevenlabs agents twilio outbound_call \
--agent-id "your-agent-id" \
--agent-phone-number-id "your-phone-number-id" \
--to-number "+1234567890" \
--call-recording-enabled true
```
See [Outbound Calls Reference](references/outbound-calls.md) for provider-specific endpoints, configuration overrides, and dynamic variables.
## Managing Agents
### Using CLI (Recommended)
```bash
# List agents and check status
elevenlabs agents list
elevenlabs agents status
# Import agents from platform to local config
elevenlabs agents pull # Import all agents
elevenlabs agents pull --agent <agent-id> # Import specific agent
# Push local changes to platform
elevenlabs agents push # Upload configurations
elevenlabs agents push --dry-run # Preview changes first
# Add tools
elevenlabs tools add-webhook "Weather API"
elevenlabs tools add-client "UI Tool"
```
### Project Structure
The CLI creates a project structure for managing agents:
```
your_project/
├── agents.json # Agent definitions
├── tools.json # Tool configurations
├── tests.json # Test configurations
├── agent_configs/ # Individual agent configs
├── tool_configs/ # Individual tool configs
└── test_configs/ # Individual test configs
```
### SDK Examples
```python
# List
agents = client.conversational_ai.agents.list()
# Get
agent = client.conversational_ai.agents.get(agent_id="your-agent-id")
# Update (partial - only include fields to change)
client.conversational_ai.agents.update(agent_id="your-agent-id", name="New Name")
client.conversational_ai.agents.update(agent_id="your-agent-id",
conversation_config={
"agent": {"prompt": {"prompt": "New instructions", "llm": "claude-sonnet-4"}}
})
# Delete
client.conversational_ai.agents.delete(agent_id="your-agent-id")
```
See [Agent Configuration](references/agent-configuration.md) for all configuration options and SDK examples.
## Error Handling
```python
try:
agent = client.conversational_ai.agents.create(...)
except Exception as e:
print(f"API error: {e}")
```
Common errors: **401** (invalid key), **404** (not found), **422** (invalid config), **429** (rate limit)
## References
- [Installation Guide](references/installation.md) - SDK setup and migration
- [Agent Configuration](references/agent-configuration.md) - All config options and CRUD examples
- [Client Tools](references/client-tools.md) - Webhook, client, and system tools
- [Using the Procedure API](references/using-procedure-api.md) - Procedure CLI and SDK flow, compile and publish
- [Writing Procedures](references/writing-procedures.md) - Trigger and content authoring, step schema
- [Widget Embedding](references/widget-embedding.md) - Website integration
- [Outbound Calls](references/outbound-calls.md) - Phone call integrations
Referenced files: 7
agent-simplification15.3 KB
---
name: agent-simplification
description: "Simplify a complex ElevenLabs voice agent (typically a large workflow agent) into a leaner architecture — fewer workflow nodes, deduplicated prompts, procedures where they help — while PROVING with a ground-truth test suite that every behaviour of the original is preserved or improved. Use when the user says things like 'simplify this agent', 'flatten this workflow', 'this agent is too complex', 'consolidate these nodes', 'rebuild this agent with fewer nodes', 'can we replace this workflow with procedures', or asks whether a big workflow agent can be made easier to maintain without losing behaviour."
---
# ElevenLabs Agent Simplification
Turn a complex agent (usually a many-node workflow) into a simpler one that is **provably at least
as good**. The core loop: understand everything the complex agent does → encode it as a test suite
(ground truth) → build the simpler agent → iterate until it matches or beats the original on every
test → A/B on branches → hand over with an honest report.
The output is never just a simpler agent. It is a simpler agent **plus the test suite that proves
equivalence**, attached natively in the platform so the client can re-run it.
## Non-negotiable safety rails
- **Never edit the production agent, global tools, the knowledge base, or shared tests.** Work on a
duplicate agent or a branch. Tools and KB documents are workspace-global — editing them changes
prod. Procedures are per-agent (cloned on duplicate), so a copy's procedures are safe to edit.
- Before any PATCH, verify you are targeting the copy/branch (`agent_id` / `branch_id` check in the
script). Make every write script refuse to run against anything else.
- Snapshot the full agent JSON before the first change (rollback + later diffing).
- **PII hygiene:** real transcripts contain names/numbers. Keep them in scratch space, never commit
them, and delete them when done. Scrub any dynamic-variable seed derived from a real call
(replace the client name, activation date, etc.). Conversation IDs are fine to reference.
## Phase 1 — Understand the complex agent completely
Pull the agent JSON (`GET /v1/convai/agents/{id}`) and map every layer:
1. **Base prompt** (`conversation_config.agent.prompt.prompt`).
2. **Workflow nodes — read `additional_prompt`, not just `prompt.prompt`.** This is the #1 trap:
an `override_agent` node's real instructions live in `additional_prompt` (the studio
"Conversation goal" box, APPENDED to the base prompt). `conversation_config.agent.prompt.prompt`
on a node is only set when the "Override prompt" toggle is ON. A node whose `prompt.prompt` is
empty is NOT an empty node. Extract every node's `additional_prompt` to a file — in a real case
16 "empty-looking" nodes carried ~100k chars of specialized business logic.
3. **Edges** — routing conditions (`forward_condition.label`), in/out degrees. The only **dead
nodes** are orphans (no incoming edge = unreachable) and fully disconnected nodes (no edges at
all). Anything with edges is a live decision point — every node has an LLM choosing among its
outgoing edges, even a tool node with an empty tool list — so it cannot be considered for
removal unless genuine logic duplication is proven.
4. **Per-node scoping** — `additional_knowledge_base` / `additional_tool_ids`. Check whether the
base agent already exposes the same KB (folders in RAG `auto` mode = every doc) and tools; if so
the scoping is redundant, not behaviour.
5. **Say nodes** — fixed messages the client wants said verbatim (goodbyes, transfer preambles,
mandated phrases). These become wording-parity tests later.
6. **Transfer machinery** — compare workflow `phone_number` node destinations against the base
`transfer_to_number` system tool (destination SIP/number + its `condition` text). Often
identical → the workflow transfer scaffolding is redundant. Check real usage: count tool fires
across recent conversations (which transfer path actually runs in prod?).
7. **Dispatch tools** (tool nodes in the graph) — these fire **deterministically** whenever the
graph reaches that node, whereas agent-attached tools fire at the LLM's discretion. Examine each
one: which paths reach it, and what guarantee does it provide? Flattening removes that guarantee,
so every dispatch path needs a test proving the tool still fires in the simplified agent. Some
dispatch tools also turn out redundant (e.g. a "say X then transfer" preamble tool duplicating
what the `transfer_to_number` system tool already does) — check real-call tool-fire counts before
deciding.
8. **Procedures, system tools (`built_in_tools`), language config** — note `language_presets` lives
at the TOP level of `conversation_config`, not under `.agent`.
9. **Real conversations** — pull recent calls (`GET /v1/convai/conversations?agent_id=&page_size=100`,
paginate; detail per call). Categorize (use the agent's own `call_category` data-collection if
present), count tool fires, and note failure modes. This tells you which paths matter and at what
frequency.
If you find duplicated logic anywhere (repeated boilerplate across nodes, per-node scoping the base
already has, scaffolding duplicating a system tool), simply remove it in the simplified build —
after confirming it is a true duplicate, not a variant.
## Phase 2 — Build the ground-truth test suite BEFORE simplifying
Derive test scenarios from **both** sources:
- **The complex agent itself**: every distinct behaviour in the node `additional_prompt`s, the base
prompt's flows (e.g. an operator-transfer escalation sequence), each qualifying-transfer case,
say-node wording, dispatch-tool paths, procedure flows, tool-consent rules, deprecated services,
coverage/country rules, language policy.
- **Real conversations**: one scenario per distinct path/category actually observed, phrased as a
simulated-user persona (reference the source `conv_` id in the scenario for traceability).
Two complementary instruments — pick per behaviour:
| | `simulation` tests (multi-turn) | `llm` tests (single-turn) |
|---|---|---|
| What | Simulated user persona converses N turns; whole conversation judged against `success_conditions` | Fixed `chat_history`; judge the agent's NEXT reply against `success_condition` |
| Use for | Procedure-driven and multi-step flows (diagnostics, escalation sequences, retention) | Sharp single decision points (don't transfer on first demand; reply language; refusal) |
| Avoid for | — | Any flow whose first agent turn is a tool call (`start_procedure`, `language_detection`): the judge sees no text → `unknown`. Use a simulation test instead. |
Criteria-writing rules (learned the hard way):
- **Verify every fact in a criterion against the KB before asserting it.** Do not label something a
hallucination until you've grepped the KB (a "made-up" USSD code turned out to be documented).
- **Accept reasonable behaviour.** If answering a direct question briefly before a still-needed
transfer is fine, say so in the criterion. Over-strict criteria create false failures you'll then
"fix" wrongly.
- **Don't let the scenario undermine the premise.** If testing "plain SIM fails where partner is
4G-only", pin a country that actually has no 3G (check the KB), or the agent will be correct and
the test wrong.
- **Don't test `end_call`.** Test harnesses don't simulate call termination reliably, and the caller
hangs up anyway. Testing that the agent does NOT hang up prematurely is fine.
- One criterion per saved test where possible (clean per-behaviour pass/fail in the dashboard).
Mechanics:
- Save as **native tests** (`POST /v1/convai/agent-testing/create`) and attach to both the
**original (reference) agent** and the **simplified copy** via
`platform_settings.testing.attached_tests` — dashboard-visible, re-runnable by the client, and
the same test objects score both sides of the A/B.
- Mock every webhook/client tool with **`tool_mock_overrides`** (per-test inline mock bodies keyed
by tool id: `{tool_id: [{"mock_result": "<json string>"}]}` with
`tool_mock_config: {mocking_strategy: "all", ...}`). Nothing hits real endpoints and no global
tool object is edited.
- Seed `dynamic_variables` from a real call (scrubbed), and include `system__call_sid`,
`system__caller_id`, `system__conversation_id`, `system__called_number` — simulations error with
`missing_dynamic_variables` otherwise if any tool references them.
- Tests already attached to the original agent are shared objects: adding new tests alongside them
is fine, but never edit or remove the existing ones.
**Baseline the original agent on the suite first.** Its failures are improvement opportunities, not
excuses ("equivalently bad" is not the goal).
## Phase 3 — Simplify
Default target: **single-agent (1-node workflow)** + base prompt + focused rules + procedures + KB.
Keep the workflow only where determinism is genuinely required and prompt enforcement proves
unreliable in tests.
- Flatten the workflow (`PATCH` with a start-node-only workflow). Check what the main/hub node's
`additional_prompt` actually adds beyond the base prompt — if it adds nothing (or duplicates it),
there is nothing to port from that node.
- Port each node's **unique** logic (not any repeated boilerplate) into either:
- **focused prompt rules** — short `#`-titled sections, one behaviour each; or
- **procedures** — for genuinely multi-step flows (diagnostics, retention). See the procedures
API lifecycle below; procedures on the copy are safe to edit (per-agent clones).
- Deterministic phrasing the client mandated (say nodes) → encode single verbatim phrases as
explicit prompt rules and test them. **Mandated multi-step sequences** (e.g. "clarify twice, ask
a feedback question, only then transfer") belong in a **procedure** — step ordering is exactly
what procedures are for, and prompt-rule enforcement of step order is less stable (the model
tends to jump to the terminal action once the customer has pushed enough times, skipping the
last mandated step). Whichever mechanism you use, test the full sequence multi-turn.
- Common config fixes while you're there: `PATCH` rejects a body containing both resolved `tools`
and `tool_ids` — pop `tools`, keep `tool_ids`. Read-modify-write the full `conversation_config`
so nothing is wiped.
## Phase 4 — Iterate on failures (the discipline that makes this work)
Run the suite on the copy; then, for **every** non-passing test (on the copy **or** the original):
1. Read the transcript/rationale once. **Do not re-run to confirm a failure** — one failure means
there's something to improve. (Re-running to verify a FIX is fine.)
2. Classify: **agent bug** (fix with a focused rule/procedure edit, then verify) vs **test bug**
(wrong premise, over-strict criterion, wrong instrument — fix the test, never by loosening it
past what's actually correct) vs **ungradable artifact** (tool-only turn in an `llm` test →
convert to a simulation test or retire the duplicate).
3. Failures shared by the original agent are still yours to fix — the tests encode ground truth
from real calls, and the goal is a better agent, not parity with its defects.
4. Watch for **rule dilution**: as prompt rules accumulate, earlier enforcement can weaken. If a
previously-passing mandated behaviour regresses, strengthen that rule's self-check rather than
adding more prose.
## Phase 5 — Final A/B on branches of the original agent
The cleanest comparison is two branches of the SAME agent:
- `POST /v1/convai/agents/{id}/branches` with `{parent_version_id, name, description,
conversation_config, platform_settings, workflow}` → a branch (e.g. "Pete") carrying the
simplified config; Main stays untouched. `parent_version_id` = the agent's current `version_id`.
- `GET /v1/convai/agents/{id}?branch_id=...` to verify each branch's config.
- Run the full suite per branch: `POST /v1/convai/agents/{id}/run-tests` with
`{"tests":[{"test_id":...}], "branch_id": ...}`; poll
`GET /v1/convai/test-invocations/{invocation_id}` until all runs are terminal; result is
`condition_result.result` (`success` / `failure` / `unknown`; `rationale` may be a dict with
`summary`).
- Report per-branch totals + per-test diffs, with invocation ids (client-verifiable in the
dashboard).
## Phase 6 — Report honestly
- State the claim precisely: "reproduces every tested behaviour of the original on a much simpler
architecture and scores X vs Y on the suite" — not "a 1-node agent is inherently better".
- Name the caveats: LLM-judged simulations (strong signal, not live traffic), single-run noise,
you authored the tests you're scoring against, per-turn token cost of a bigger always-on prompt,
and workflow determinism vs prompt enforcement for verbatim phrases.
- Recommend human review of a few calls and a gradual rollout before any full switch.
- If you got something wrong along the way, correct the record explicitly in the report.
## API gotchas appendix
**Procedures (draft → commit lifecycle — everything is branch-scoped):**
- `POST /v1/convai/agents/{id}/branches/{br}/procedures` `{name, content, type: "free_form",
trigger?}` → returns `procedure_id`, but creates the procedure **as a draft** and adds its ref to
your user's agent draft. It will NOT appear in the committed agent/branch procedure lists yet —
do not conclude the API "didn't attach" (checking the committed view after create is the classic
mistake).
- `GET .../procedures/status` → the draft refs; `GET agent?include_draft=true` → draft config.
- `POST .../procedures/compile` → **preview only**: validates and returns the compiled workflow +
errors. It commits nothing.
- **Commit = `PATCH /v1/convai/agents/{id}?branch_id={br}`** (an empty JSON body works): it
resolves procedure refs from your draft into the new committed version, then deletes the drafts.
After this, the procedure appears in `GET agent` `.procedures`, the branch list, and direct GET.
- Edit content: `PATCH .../procedures/{pid}/draft` (full body: name, content, type, trigger) →
commit PATCH. Delete: `DELETE .../procedures/{pid}` (removes from draft set) → commit PATCH.
- Procedure `content` = markdown with `---` frontmatter (`name`, `trigger`) + steps. Reference KB
docs as `[kb id="..." name="..."]`, tools as `[tool id="..." name="..."]`, system tools as
`[system_tool id="transfer_to_number" name="..."]`.
- Procedures are per-agent (cloned when an agent is duplicated) — but the KB/tools they reference
are global.
**Testing:**
- Saved-test update is `PUT` (full replace — a partial body silently wipes omitted fields), not
`PATCH` (405).
- `simulate-conversation` (ad-hoc, no saved artifact) runs the LIVE agent config — to A/B a
candidate config you must actually PATCH it (or use branches + `run-tests`).
- A test whose judged turn is a bare tool call grades `unknown` — instrument mismatch, not failure.
- Multi-turn sims have some single-run flakiness. Don't paper over it with repeat runs — testing
has cost; prefer broadening the suite with more varied scenarios, and investigate any failure
rather than re-rolling it.
**Agent config:**
- A language preset makes a language switchable by `language_detection`; but if a `{{language}}`
dynamic variable anchors the default, add an explicit "the spoken language wins" prompt rule.
- Conversation-detail endpoint rate-limits hard — keep concurrency ~4-5 with retries.
agents-platform2.92 KB
--- name: agents-platform description: Manage ElevenLabs voice agents through the connected ElevenLabs MCP server. Use when the user asks to create, configure, test, deploy, or inspect agents, knowledge bases, agent tools, tests, or conversations directly on their ElevenLabs account — rather than writing application code. Requires the ElevenLabs MCP server bundled with this plugin. license: MIT --- # ElevenLabs Agents Platform (via MCP) Workflow guidance for the `agents_*` tools on the ElevenLabs MCP server. Every tool takes a `context` parameter — briefly state the user's goal in it. For building agents in the user's own codebase (SDK, CLI, widget embedding), use the general `agents` skill instead. ## Creating and updating agents - Create with `agents_create` using the top-level fields (`name`, `prompt`, `first_message`, `voice_id`, `language`, `tags`). Only reach for the raw `body` escape hatch when a field has no top-level equivalent. `voice_id` is optional — the workspace default voice is used if omitted. - Before changing an existing agent, read it: `agents_list` to find it, `agents_get` to see its current config, then `agents_update` with only the fields that should change. - `agents_duplicate` clones an existing agent — prefer it over recreating a config by hand. - Get a shareable/test link with `agents_get_link`; inspect widget config with `agents_get_widget`. ## Knowledge base - Add documents with `agents_create_kb_url` (supports auto-sync for pages that change) or `agents_create_kb_text`; organize with `agents_create_kb_folder`. - Before deleting a document, check `agents_get_kb_dependents` — other agents may use it. Confirm with the user before `agents_bulk_delete_knowledge_base`. - Verify retrieval quality with `agents_query_knowledge_base_rag` or `agents_search_knowledge_base` after adding documents. ## Testing and deployment - Create tests with `agents_create_test`: an LLM response test (`success_condition` prompt plus success/failure examples), a tool-call test (verify a tool was or wasn't called with expected parameters), or a simulation test (`simulation_scenario` persona plus `success_conditions`). - Run and inspect with `agents_list_test_runs` / `agents_get_test_run`. - For staged changes, branch: `agents_create_branch`, make edits, preview with `agents_merge_branch_preview`, then `agents_merge_branch` and `agents_create_deployment`. ## Conversations - `agents_list_conversations` and `agents_get_conversation` inspect call history and transcripts; `agents_search_conversation_messages` finds specific exchanges. Use these to debug agent behavior before editing prompts. ## Safety - Confirm with the user before destructive calls: `agents_delete`, `agents_delete_tool`, `agents_delete_phone_number`, knowledge-base bulk deletes. Deletions are not reversible. - IDs (agent, document, tool, branch) only ever come from tool results — never guess or reconstruct them.
architect-agent-memory3.18 KB
---
name: architect-agent-memory
description: Use when the user wants the agent to remember, learn, save, or memorize a fact, or reports a one-off "the agent said X but should say Y" correction. There is no memory endpoint over REST; this covers the durable alternatives (system prompt or knowledge base) and when to use them.
---
# Teach an agent a fact (there is no memory endpoint over REST)
The web Architect has a `create_memory_entry` action for ad-hoc, agent-global facts. Over the ConvAI REST API there is NO memory endpoint. Say this plainly to the user rather than implying you saved a memory. Instead, persist the fact one of two durable, branch-scoped ways below (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`; you need `$AGENT_ID` and `$BRANCH_ID`).
## First, sanity-check the request
- Confirm the correction is scoped and factual: a discrete fact or rule ("In v3, only the Stability slider is supported"), not a broad behavioral change. Broad, recurring, or procedural changes belong in the system prompt or a new procedure, not a one-off fact.
- Confirm the fact genuinely cannot come from a crawled knowledge base (which can re-sync) or an API call (which stays current). If it can, prefer those sources.
- If the request is a broad behavioral change, escalate to editing the system prompt or creating a procedure instead.
## Option A: a durable fact in the system prompt
Add the fact as a short, self-contained line the agent can recall verbatim. Read the current prompt scoped to the branch, append the line, and PATCH the whole string back:
- Read: `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`, take `conversation_config.agent.prompt.prompt`.
- Write: `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` with `{"conversation_config":{"agent":{"prompt":{"prompt":"<full edited prompt>"}}}}`.
This commits to the branch HEAD and returns a new version. Keep the fact short so it does not bloat the prompt.
## Option B: a knowledge base document
For a fact better kept as recallable knowledge than a prompt line, create a KB text doc and attach it to the agent:
- `POST /v1/convai/knowledge-base/text` with the fact as the document body.
- Attach the returned document to the agent (add its id to the agent's knowledge base config via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`).
## Timing matters: both options are branch-scoped
Unlike the web memory feature (which is agent-global and goes live immediately), both the prompt line and the KB doc are branch-scoped. They only reach the production agent when the branch is merged. This is usually what you want: it lets the user stage the fact for a launch or embargo and merge at go-live. Tell the user the change is staged on the branch and ships on merge.
If the user explicitly wants a fact live on production right now and is comfortable with that, note that the durable options here still require a merge; there is no REST path to an immediate agent-global memory. Make that limitation explicit rather than pretending otherwise.
## Confirm
After writing, briefly confirm to the user what was saved, where (prompt line or KB doc), on which branch, and that it goes live on merge. Keep it to one or two short sentences.
architect-branches-versions-merge6.48 KB
---
name: architect-branches-versions-merge
description: Use when working with agent branches, versions, drafts, traffic splits, or merges. Fires on "make a copy to test on", "create a branch", "switch to my test branch", "send 10% of traffic to the new version", "roll it out gradually", "merge my branch back to main", or when a branch operation errored.
---
# Branches, versions, and merging
Branches are versioned snapshots of an agent's config (prompts, voices, tools, workflows) that let you test changes without touching live production. Main is the default branch and receives 100% of traffic unless a split is configured. Drafts are work-in-progress edits on a branch that are not committed to a version until published. Traffic splitting distributes a share of live conversations deterministically across branches (by conversation id) for A/B testing. Merging folds a source branch's changes into its parent (usually main), creating a new version on the target.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY` and `$AGENT_ID`.
Branch operations fail most often on stale state: a wrong branch id, a traffic allocation that does not total 100, or an operation blocked by branch protection. Read current branch state first, then route to the exact operation.
## HARD GATE: resolve the id before any write
Never write to a branch id you have not just confirmed. If the user names a branch ("switch to angelo/test-switch", "go to my test branch"), that name is not a ready-to-use id. Resolve it by listing branches and matching on `name` first, in the same task, before the first write. Do not reuse an id from earlier context, memory, or a guess. If no exact name match exists, say so and ask the user to confirm rather than picking the closest-looking one.
## 1. Read current branch state first
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches" \
-H "xi-api-key: $API_KEY"
```
This is the source of truth. For each branch note the exact `id` (agtbrch_...), `current_live_percentage`, `protection_status` (admin_perms_required vs writer_perms_required), `parent_branch_id`, `main_branch_id`, and `draft_exists`. Get one branch's true HEAD version from `.most_recent_versions[0].id`.
Also get the agent to confirm which branch is currently active and the agent id:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID" -H "xi-api-key: $API_KEY"
```
## 2. Route to the exact operation
- **Create a copy to experiment on** -> `POST /v1/convai/agents/$AGENT_ID/branches` with `{name, description, parent_version_id}`. Use the parent branch's HEAD version id (from `.most_recent_versions[0].id`). New branches start at 0% live traffic, so creation alone sends no live traffic.
- **Edit config on a different branch** -> scope every read and write to that branch with `?branch_id=<id>`. `PATCH /v1/convai/agents/$AGENT_ID?branch_id=<id>` commits to that branch's HEAD, never main.
- **Send a share of live traffic to a branch** -> set the traffic split. This is the operation most likely to fail; see the constraints below.
- **Fold a branch's changes back into its parent** -> merge. Preview first, then merge; the merge is destructive on the target.
## 3. Constraints that cause the failures
1. **Traffic split must total exactly 100% across all active branches.** Setting a split declares the whole allocation, not one branch. If you raise a new branch to 10%, lower another (usually main) by the same 10% in the same call. Sum every branch's `current_live_percentage` from the branch list, apply the delta, and confirm the new set sums to 100 before sending. A partial set that does not total 100 is the top validation failure.
2. **Only include active (non-archived) branches** in the split. An archived branch or a stale id triggers not_found or validation. Rebuild the allocation strictly from the current branch list.
3. **Protection status blocks writes.** Main typically has `protection_status: admin_perms_required`. If the user is not an admin, merging into main and editing main fail with a permission error. Check `protection_status` before promising a merge; if it is admin-gated and the user lacks the role, route them to the traffic-split A/B path (ramp traffic instead of merging) or to support.
4. **Merge target follows parentage.** A branch merges into its `parent_branch_id`. Verify the branch has the parent the user expects. You cannot merge a branch into an unrelated branch.
5. **An uncommitted draft blocks the operation.** If `draft_exists: true`, unsaved edits can cause a wrong-state error on switch or merge. Have the user save/commit or explicitly discard the draft first. Do not silently overwrite it.
6. **Ramp gradually, do not jump to 100.** Move a proven branch 10% -> 50% -> 100% across separate traffic-split calls, watching analytics between steps, rather than one 0 -> 100 jump.
7. **Traffic split is not the same as explicit branch selection.** `current_live_percentage` only governs how un-pinned live traffic is auto-routed. A branch at 0% can still receive conversations when it is selected explicitly (a `branch_id` passed via the API, or an explicit branch pin in a client). Never tell a user that a 0% branch gets no conversations, and account for explicitly-targeted branches when reasoning about where recent conversations came from.
## 4. Merge
Merge folds a branch into its parent. Always run a merge preview first to show the diff before merging. The merge endpoint is:
```
POST /v1/convai/agents/{agent_id}/branches/{source_branch_id}/merge
```
Merging is destructive on the target and creates a new version on it. If main is `admin_perms_required` and the user lacks the role, the merge fails with a permission error; do not retry. Offer the traffic-split A/B path instead, or escalate to support.
## 5. Recovery
- A bare validation failure is almost always the traffic split not summing to 100 or an array shape issue. Re-read the branch list, recompute the full allocation over active branches to total exactly 100, resend. Do not retry the identical payload.
- not_found means the branch id is stale, archived, or a name was passed. Re-read the branch list, copy the exact current id, retry. For a traffic split, drop any branch not present in the fresh list.
- A wrong-state error means you acted before the branch was ready or an uncommitted draft is blocking. Save/commit or discard the draft, then retry.
- A permission error means the target branch is protection-gated and the user lacks the role. Do not retry; offer the traffic-split path or escalate.
architect-create-client-tool5.42 KB
---
name: architect-create-client-tool
description: Use when the user wants the agent to trigger something in their own browser, app, or SDK during a call (navigate a page, open or close a dialog, read on-screen context, run client-side code), or asks "how do I add a client tool", "make the agent do something in my app", or reports a client tool that won't save or where the agent never gets a response back. For a server API tool see the create-webhook-tool skill; for editing an existing tool see the edit-existing-tool skill.
---
A client tool runs in the caller's browser or app via the widget or SDK, not on ElevenLabs' servers. The agent emits a call over the event stream and the customer's own code executes it. This is different from a webhook tool (ElevenLabs calls an HTTP endpoint). All API calls use `xi-api-key: $API_KEY` against `https://api.elevenlabs.io`.
## 1. Gather current state first
Before creating, make these reads (in parallel where independent):
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` to confirm the agent, its LLM, and whether it is a workflow agent (workflow agents scope tools per node via `additional_tool_ids`, so where you attach the new tool matters; the workflow is in this config).
- `GET /v1/convai/tools` to avoid creating a duplicate and to reuse an existing client tool if one fits.
## 2. Create the tool
```
POST /v1/convai/tools
```
with the tool config discriminated by the client type. The fields that cause failures:
- `name`: stable snake_case identifier the SDK matches on. Keep it identical to what the client code listens for; renaming later breaks the integration.
- `description`: the LLM uses this to decide WHEN to call the tool. Be specific ("Navigate the user to the billing settings page", not "navigation"). Vague descriptions are the top cause of "the agent never calls it".
- `parameters`: an ARRAY of parameter objects (not a JSON-Schema object with a `properties` map). Each item needs: `id` (snake_case, non-empty), `type` as a BARE string (`"string"`|`"number"`|`"integer"`|`"boolean"`; nullable uses a 2-element array like `["string","null"]`, never `{"type":"string"}`), `description`, `value_type` (`"llm_prompt"`|`"dynamic_variable"`|`"constant"`), `dynamic_variable` (empty string unless bound), `constant_value` (always a string/number/boolean, NEVER null or omitted), and `required` (boolean). Exactly one value source must be populated: a non-empty `description` (llm_prompt), OR `dynamic_variable`, OR a non-empty `constant_value`, OR `is_system_provided: true`. Zero sources is invalid and is the top cause of `invalid_union` / `schema_mismatch` on `parameters.0`. A working llm_prompt item: `{"id":"city","type":"string","description":"The city to look up.","value_type":"llm_prompt","dynamic_variable":"","constant_value":"","required":true}`. For no inputs send `parameters: []`. Array-typed params also need an `items` schema.
- `expects_response`: the single most important decision.
- `true`: the agent PAUSES and waits for the client to send a result back (bounded by `response_timeout_secs`, default 20, max 120). Use when the return value feeds the conversation ("read the current page and tell me what the user is looking at").
- `false`: fire-and-forget; the agent does not wait. Use for pure side effects (navigate, close dialog). Setting `true` for a side-effect-only tool is a common mistake; the agent then hangs until timeout because the client never sends a result.
- `execution_mode`: `immediate` runs right away; `post_tool_speech` lets the agent speak first. Async / fire-and-forget maps to `expects_response: false`. Only set `response_timeout_secs` when `expects_response: true`, and keep it <= 120.
### Client event contract (tell the user; this is where their end breaks)
The tool only works if the client handles the call. On invocation the client receives a `ClientToolCall` event carrying `tool_call_id`, `tool_name`, and `parameters`. If `expects_response: true`, the client MUST send back a `ClientToolResult` with the SAME `tool_call_id` and a string result within `response_timeout_secs`, or the agent gets a timeout. So `expects_response: true` with no client handler = every call times out; that is a client-code gap, not a tool-config bug.
## 3. Attach and instruct
Attach the tool to the agent via a targeted `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`: on a single-node agent add its id to `tool_ids`; on a workflow agent keep base `tool_ids: []` and add it to the right node's `additional_tool_ids` (see the edit-workflows skill for node scoping). In the prompt, tell the agent when to call it and what to do on failure ("If it returns an error or times out, tell the user you could not complete the action").
## 4. Recovery
- `schema_mismatch`: usually the `parameters` block. Confirm each item has exactly one value source and a bare-string `type`, and that array params have an `items` schema. Rebuild minimally and retry.
- `validation`: a required field is missing or malformed (empty `name`, missing `description`, `response_timeout_secs > 120`). Fill or clamp and retry.
- `not_found`: a wrong agent, tool, or node id. Re-read current ids and route to the correct one.
## 5. Verify
After creating and attaching, re-`GET /v1/convai/tools` (and the agent config for a workflow) to confirm the tool exists and is scoped to the intended node, then confirm with the user that their client code listens for `ClientToolCall` and returns a `ClientToolResult` when `expects_response` is true.
architect-create-llm-test5.03 KB
---
name: architect-create-llm-test
description: Use when the user wants a single-turn test that checks WHAT the agent says for a given turn. Fires on "add a test that the agent greets the caller", "test that it always reads the disclosure", "make sure it refuses off-topic questions", "write an eval for this reply", or "add a pass/fail check on the agent's answer". Not a multi-turn simulation and not a tool-call check.
---
# Create an LLM response test
Use this when the user wants to assert that the agent's reply to a turn matches a natural-language success condition. This is a single-turn response judgement. If they instead want to verify which tool fires, use the tool-test skill; if they want a multi-turn role-played conversation, use the simulation-test skill. For a DOM or Architect-style agent whose first action is almost always starting a procedure, prefer a simulation test instead, since a single-turn LLM test only grades that first action and returns useless verdicts.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID` where relevant.
## 1. Gather context first
Read only what you need to ground the test, in parallel where independent:
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` - to ground the success condition in the agent's actual behavior: its first message, system prompt, and any required disclosures. A condition like "greets the user" is worthless if you do not know what the greeting is supposed to be.
- List existing tests to avoid a near-duplicate and to copy the naming convention already in use.
If the user is turning a real bad call into a regression test, read the conversation (`GET /v1/convai/conversations?agent_id=$AGENT_ID`, then `GET /v1/convai/conversations/{conversation_id}`) and lift the transcript into `chat_history`. Conversation reads carry customer PII; do not copy them into other systems or logs, and respect zero-retention-mode accounts.
## 2. Create the test
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agent-testing/create" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{
"type": "llm",
"name": "Greets caller by name on open",
"success_condition": "The reply welcomes the caller and offers help. Mark SUCCESS only if it does both.",
"chat_history": [{"role": "user", "message": "Hi"}]
}'
```
An LLM test needs exactly three fields:
- `name` - non-empty string. Reuse the naming style from the existing tests.
- `success_condition` - natural-language description of what makes the reply correct.
- `chat_history` - array of `{role, message}` where `role` is `"user"` or `"agent"` only, and the array must end on the `user` turn the agent is supposed to respond to.
## 3. Schema gotchas that cause failures
1. **Empty or missing `chat_history`.** It is required and must be non-empty, ending on a `user` turn. An empty array fails with "Chat history is required for llm tests".
2. **Wrong `role` value.** The enum is strictly `"user"` or `"agent"`. Do not use `"assistant"`, `"system"`, `"bot"`, or `"human"`.
3. **Blank `success_condition`.** A whitespace-only string also fails. Write a concrete criterion.
4. **Vague `success_condition`.** It validates but produces a flaky test. Phrase it as an evaluation question with an explicit pass bar ("Mark SUCCESS only if the agent reads the disclosure starting with X"), not as an instruction ("The agent should read the disclosure").
5. **Wrong test type for the intent.** If the user wants to check a tool call, `success_condition` has no effect; use the tool-test skill. If they want a multi-turn role-play, use the simulation-test skill (`chat_history` is not used there). Sending simulation fields to an LLM test is a common schema mismatch.
6. **Putting the expected answer in `chat_history`.** The final turn must be the user prompt. Do not add a trailing `agent` message with the "right" answer; the agent generates that reply at run time and the judge scores it against `success_condition`.
7. **`chat_history` over 200 messages is rejected.** Trim long transcripts to the turns that set up the one under test.
## 4. Attach and run
Creating a test only registers it. If it should run going forward, attach it to the agent and branch, then run it:
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/testing/attach-test" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"test_id": "'"$TEST_ID"'", "branch_id": "'"$BRANCH_ID"'"}'
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"tests": [{"test_id": "'"$TEST_ID"'"}], "branch_id": "'"$BRANCH_ID"'", "repeat_count": 1}'
```
Poll `GET /v1/convai/test-invocations/{suite_id}` until `test_runs[0].status` leaves `pending`, then read `condition_result`. If the user wants several checks, create them one at a time. If a created test keeps failing on a valid conversation, the `success_condition` is likely too strict; loosen it or split compound criteria.
architect-create-simulation-test7.36 KB
---
name: architect-create-simulation-test
description: Use when the user wants a full multi-turn conversation test where a simulated user talks to the agent across many turns. Fires on "test the whole flow", "simulate a caller", "write a scenario test", "test that the agent handles an upset customer / a full booking / a multi-step call", or when a simulation-test creation attempt failed. For single-response checks use the LLM-test skill; for one tool call use the tool-test skill.
---
# Create a simulation test
A simulation test drives a full multi-turn conversation: a simulated user (defined by a persona and scenario) talks to the agent, then the transcript is judged against success conditions. It is the highest-value test type and empirically the most failure-prone, almost always because of a schema or wrong-type mistake, not a bad idea. This skill gets it right first time.
**Always use a simulation test for DOM or Architect-style agents.** Their first action is almost always starting a procedure, so LLM and tool tests grade only that first action and return useless verdicts. Only the multi-turn simulation exercises the real flow.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID` where relevant.
## 1. Read the agent first
So the persona and success conditions match reality, read these in parallel where independent:
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` - the system prompt, first message, language, and LLM. The persona and success conditions must reflect what the agent is actually instructed to do. Confirm the agent language so the simulated-user prose is in a language the agent handles. If the agent has a workflow (in this same config), write the scenario to walk the simulated user through the node transitions you want exercised.
- List the agent's tools for the ids you need in `tool_mock_config` and to name a specific tool in a success condition.
- List existing tests to extend coverage instead of duplicating, and read a nearby existing simulation test to copy its exact field shape. Mirroring a working example is the single most reliable way to avoid a schema mismatch.
## 2. Create the test
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agent-testing/create" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{
"type": "simulation",
"name": "Upset customer, surprise charge, de-escalate then retain",
"simulation_scenario": "You are Diane, a loyal customer upset about a surprise $45 charge. Acknowledge the greeting, ask why you were charged, then ask for a refund. Once resolved, say you are done.",
"success_conditions": ["The agent verified the account before discussing the charge", "The agent de-escalated before quoting policy", "The agent offered a retention option"],
"simulation_max_turns": 20,
"tool_mock_config": {"mocking_strategy": "all", "fallback_strategy": "raise_error", "mocked_tool_ids": []},
"dynamic_variables": {"tier": "enterprise"},
"chat_history": []
}'
```
Keep the two prose fields distinct:
- `simulation_scenario` - instructions to the simulated user (the persona plus what they do, in second person). This drives the other side of the conversation. Give the simulated user a clear stop cue so the conversation ends.
- `success_conditions` - plain strings, the pass/fail rubric judged against the agent's behavior, written as evaluation statements, not instructions. Enumerate concrete checkable clauses. Reference a tool by its real name if success requires a tool call. Vague conditions make the test flaky.
Other fields:
- `type` must be exactly `"simulation"`.
- `simulation_max_turns` - cap the conversation (existing tests use around 20-24). Too low and the flow cannot complete; too high wastes runs. For a DOM or Architect-style agent set it to at least 2, or the run comes back inconclusive.
- `tool_mock_config` - an object, not a list. Shape: `{"mocking_strategy": "all", "fallback_strategy": "raise_error", "mocked_tool_ids": []}`.
- `dynamic_variables` - a flat string-to-string object of any placeholders the agent's prompt or first message needs. Missing a required placeholder makes the sim behave unexpectedly; provide every placeholder the agent references.
- `chat_history` - leave `[]` for a fresh conversation; only prefill turns when testing behavior that assumes prior context. Each prefilled entry needs `time_in_call_secs`. A DOM or Architect-style agent usually needs a seeded `chat_history` plus mocking or every run times out.
## 3. Mocking strategy and its caveats
`mocking_strategy` is exactly one of `all`, `selected`, or `none`. There is no fourth value.
- `all` mocks every tool call with the generic fallback and ignores any per-tool overrides. Combined with `fallback_strategy: "raise_error"` and no overrides, every tool call errors mid-conversation. That is fine only when the behavior under test does not depend on a tool succeeding (e.g. a confirmation-wording fix).
- If the fix depends on a specific tool returning realistic data (e.g. a list call returning a branch list so the agent can resolve a name to an id), you must use `selected` with `mocked_tool_ids: [...]` naming every tool to mock, and provide the overrides. Under `all`, overrides silently no-op and you see "no mock matched" in the transcript even though the config saved correctly.
- Any id in `mocked_tool_ids` must be a real tool id. Mock by id, judge by name.
- Some SYSTEM tools such as RAG cannot be mocked and still run live; account for that when a run behaves unexpectedly.
## 4. Editing a simulation test is delete and recreate
There is no in-place edit for a simulation test; a PUT is response/llm only and rejects `type: simulation`. If a run reveals the config was wrong, `DELETE /v1/convai/agent-testing/{test_id}`, fix the payload, and create again, then re-attach and re-run.
## 5. Attach, then run
Creating a test only registers it; it does not add it to the agent's suite. If it should run going forward, attach it in the same task without waiting to be asked, then run:
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/testing/attach-test" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"test_id": "'"$TEST_ID"'", "branch_id": "'"$BRANCH_ID"'"}'
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"tests": [{"test_id": "'"$TEST_ID"'"}], "branch_id": "'"$BRANCH_ID"'", "repeat_count": 1}'
```
Simulation runs take minutes. Poll `GET /v1/convai/test-invocations/{suite_id}` until `test_runs[0].status` leaves `pending`, using a background poll loop rather than a foreground sleep. Then read `condition_result.result` and the per-criterion rationale, and iterate the persona or success conditions based on where the run diverged.
## 6. Recovery per error
- A schema mismatch usually means `tool_mock_config` was sent as a list instead of an object, `type` was not exactly `"simulation"`, or `success_conditions` was missing. Fetch a working simulation test and mirror its exact keys.
- A validation error means an unknown `mocked_tool_ids` id, an out-of-range `simulation_max_turns`, or a non-string `dynamic_variables` value. Fix the offending value.
- not_found means a referenced agent, tool id, or test id does not exist on this branch. Re-read the current ids and confirm the branch.
architect-create-tool-test6.76 KB
---
name: architect-create-tool-test
description: Use when the user wants a test asserting the agent CALLS a specific tool (with specific parameters, or does NOT call it). Fires on "test that it calls my booking tool", "verify it passes the right customer_id", "make sure it never calls transfer here", "add a tool test", "check it uses save_result at the end", or when a tool-test creation attempt errored.
---
# Create a tool-call test
This is the reliability drill-down for the tool-call test type. It fails often, almost always on the `tool_call_parameters` shape (wrong `eval` type, wrong `referenced_tool`, bracket paths). A tool test replays your `chat_history` turns, then checks whether the agent's next action is the expected tool call. For a single agent reply use the LLM-test skill; for a multi-turn role-play use the simulation-test skill.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID` where relevant.
## 1. Gather context first
You cannot write a correct `referenced_tool` without the target tool's real `id` and `type`, and you cannot write a realistic `chat_history` without knowing what the agent is supposed to do. Read these first, in parallel where independent:
- List the agent's tools (`GET /v1/convai/tools`, or the tool ids from `conversation_config.agent.prompt.tool_ids` on the agent config, then `GET /v1/convai/tools/{tool_id}`). Required: this is where you get the target tool's `id` and `type` (`webhook`, `client`, `code`, or `system`). `referenced_tool` needs both, and the type must match the tool's real executor type or creation fails.
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` - the prompt and first message, so the `chat_history` is a plausible lead-up to the tool call, and the base `tool_ids` so you know the tool is available to the agent. For a workflow agent, the workflow in this config tells you whether the tool is scoped to a node via `additional_tool_ids`.
- List existing tests to reuse the naming scheme and avoid a duplicate.
If grounding the test in a real call, read that conversation (`GET /v1/convai/conversations/{conversation_id}`) to lift the exact tool name and parameters the agent actually used. Conversation reads carry customer PII; do not copy them elsewhere and respect zero-retention-mode accounts.
## 2. Create the test
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agent-testing/create" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{
"type": "tool",
"name": "Saves result via save_coaching_result at end of feedback",
"chat_history": [
{"role": "agent", "message": "Great work today. Let me save this so it shows on your dashboard."},
{"role": "user", "message": "Perfect, go ahead and save it."}
],
"tool_call_parameters": {
"referenced_tool": {"id": "tool_5801k...", "type": "client"},
"parameters": [
{"path": "scenario_id", "eval": {"type": "anything"}},
{"path": "overall_score", "eval": {"type": "exact", "expected_value": "6"}}
],
"verify_absence": false
}
}'
```
## 3. Schema gotchas and how to avoid each
1. **`chat_history` role values and shape.** Each turn is `{role, message}` where `role` is exactly `"user"` or `"agent"`. The history must end on the turn just before the expected tool call, usually a `user` turn that should trigger it. An empty or agent-final history makes the assertion meaningless.
2. **`referenced_tool` needs the real `id` and `type` together.** `type` is one of `webhook` / `client` / `code` / `system` and must match the actual tool. A made-up id, or the right id with the wrong type, fails as not_found or validation.
3. **`parameters` is an array of `{path, eval}`, not a flat key/value object.** A common mistake is `parameters: {customer_id: "123"}`. The correct form is `parameters: [{"path": "customer_id", "eval": {"type": "exact", "expected_value": "123"}}]`.
4. **`eval.type` is a small enum; use the right one with its companion field:**
- `anything` - parameter must be present, value unconstrained. No extra field.
- `exact` - requires `expected_value` (a string; stringify numbers and bools, e.g. `"6"`, `"true"`).
- `regex` - requires `pattern`.
- `llm` - requires `description` (natural-language criterion the judge applies).
There is no `contains` or `semantic` type. An unknown type, or omitting the companion field, fails as a schema mismatch.
5. **`path` uses dot notation for nested args.** `path: "customer.id"` or `path: "items.0.sku"`, not `customer[id]` or `items[0].sku`. Bracket notation does not resolve.
6. **Assert only the parameters you care about.** List just the args the test should pin; leave the rest unmentioned. To assert "some tool is called, do not care which", set `check_any_tool_matches: true` and omit `referenced_tool`.
7. **`verify_absence: true` asserts the agent must NOT call the tool** given that history. Do not also fill `parameters` in that case; there is no call to inspect. Use it for "it should never transfer here" guards.
8. **`workflow_node_transition` is workflow-only.** Leave it null or omitted for single-node agents; only set it when you gathered the node structure from the workflow and want to assert the call happens after a specific transition.
## 4. Recovery per error
- A schema mismatch is almost always the `eval` object (unknown `type`, or `exact`/`regex` missing `expected_value`/`pattern`) or `parameters` passed as an object instead of an array. Fix that one field and resend.
- A validation error is usually `chat_history` (wrong `role`, empty, or not ending on the triggering user turn) or a `path` in bracket notation. Fix per gotchas 1 and 5.
- not_found means the `referenced_tool.id` does not exist on this agent or its `type` is wrong. Re-list the tools and copy the exact id and type. If the tool genuinely does not exist, create it first (see the tool-creation skills), then reference it.
## 5. Attach, run, and edit later
Creating a test only registers it. If it should run going forward, attach it to the agent and branch, then run:
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/testing/attach-test" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"test_id": "'"$TEST_ID"'", "branch_id": "'"$BRANCH_ID"'"}'
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"tests": [{"test_id": "'"$TEST_ID"'"}], "branch_id": "'"$BRANCH_ID"'", "repeat_count": 1}'
```
Pass `repeat_count` (2-50) to check flakiness. Poll `GET /v1/convai/test-invocations/{suite_id}` until the run leaves `pending`, then read `condition_result`. To edit a tool test later, read its current shape first, then update it with the full tool body.
architect-create-webhook-tool5.44 KB
---
name: architect-create-webhook-tool
description: Use when the user wants to give an agent a NEW webhook / server / HTTP tool so it can call an external API (look up a customer, create a ticket, check inventory, hit their backend). Fires on "add a webhook tool", "connect my agent to my API", "make a tool that calls my endpoint", or when a webhook-tool create fails with schema_mismatch / validation. For editing an existing tool see the edit-existing-tool skill; for a tool that saves but misbehaves at runtime see the troubleshoot-tool-errors skill.
---
Creating a webhook tool is the most common way to connect an agent to an external API, and it fails a lot, almost all from schema mistakes in `api_schema.request_body_schema`. This skill gets it right on the first write.
All calls use `xi-api-key: $API_KEY` against `https://api.elevenlabs.io`.
## 1. Gather current state first
Before writing anything, make these reads (in parallel where independent):
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` to confirm the agent and whether it is a single-node agent or a workflow (the workflow is in this config; it changes where you attach the tool later).
- `GET /v1/convai/tools` to check the tool does not already exist and to match the naming style already in use.
Get the endpoint URL, method, auth method, and request/response shape from the user before you construct the tool. If any are missing, ask for them or the API docs first.
## 2. Create the tool
```
POST /v1/convai/tools
```
with the tool config discriminated by type (webhook/server). The webhook config lives under `api_schema`, not at the top level.
## 3. Schema gotchas (the specific things that cause failures)
1. STATIC URL. `api_schema.url` must be a static string. Do NOT put `{{variables}}` in the path. Dynamic values go into `request_body_schema` (or `path_params_schema` / `query_params_schema`), never interpolated into the URL.
2. EXACTLY ONE value source per property. Every property in `request_body_schema.properties` gets its value from exactly one of: a non-empty `description` (LLM fills it), `dynamic_variable` (e.g. `"user_phone_number"` or a system var like `"system__conversation_id"`), `constant_value` (a fixed literal), or `is_system_provided: true`. Two sources on one property is the most common validation failure; ZERO sources is equally invalid. A bare `{"type":"string"}` or a property whose only non-type field is an empty `constant_value` has zero sources and is rejected. `constant_value: ""` does NOT count as a constant.
3. path_params_schema / query_params_schema are ARRAYS of parameter objects, NOT a `{properties:{...}}` object. Each item needs: `id`, `type` as a BARE string (`"string"`|`"number"`|`"integer"`|`"boolean"`; nullable uses a 2-element array like `["string","null"]`, never `{"type":"string"}`), `description`, `value_type` (`"llm_prompt"`|`"dynamic_variable"`|`"constant"`), `dynamic_variable` (empty string unless bound), `constant_value` (always a string/number/boolean, NEVER null or omitted), and `required` (boolean). A working llm_prompt item: `{"id":"city","type":"string","description":"The city to look up.","value_type":"llm_prompt","dynamic_variable":"","constant_value":"","required":true}`.
4. ARRAY properties REQUIRE an `items` schema and MUST NOT carry `constant_value`.
5. `required` must reference only property names that exist in `properties`.
6. `request_headers` is an OBJECT, not an array. For secret API keys, reference a workspace secret by id (`{"secret_id": "..."}`, list existing via `GET /v1/convai/secrets`) rather than pasting the key.
7. `response_timeout_secs` defaults to 20, max 120. Raise it only for genuinely slow APIs; high timeouts hurt voice latency.
8. Name and description drive INVOCATION. Make the description specific ("Look up a customer by phone number", not "Customer lookup"); a vague description is the top reason a correctly-saved tool never gets called.
## 4. Extract response values into the conversation
To feed API response fields into `{{dynamic_variables}}`, add entries to the tool's top-level `assignments` array (a sibling of `api_schema`). Each maps a response field to a variable: source `response`, a dot-notation `value_path` (`data.items.0.id`, NOT `data.items[0].id`, bracket indexing silently misses), and the target `dynamic_variable`. Then reference `{{customer_name}}` in the prompt so the agent speaks it.
## 5. Attach and wire usage
After creating, attach the tool to the agent. On a single-node agent add its id to the base `tool_ids`; on a workflow agent keep base `tool_ids: []` and add it to the right node's `additional_tool_ids` so it is only in scope where it fires. Apply the attachment via a targeted `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` on the relevant path. In the prompt, tell the agent exactly when to call the tool and add a fallback line ("If the lookup returns an error or no results, apologize and offer to take a message").
## 6. Recovery
- `schema_mismatch` (often no field detail): re-check the section-3 rules one by one, most often two value sources on one property, a `{{var}}` in the URL, or `constant_value` on an array. Fix and resend; do not blind-retry the identical payload.
- `validation`: a concrete arg is wrong (a `required` name with no matching property, an array missing `items`, a bad `value_path`, a non-object `request_headers`). The message names the field; correct that one.
- `not_found`: a bad agent id or a workflow node id that does not exist. Re-read current ids and retry.
architect-edit-existing-tool4.61 KB
---
name: architect-edit-existing-tool
description: Use when the user wants to change an existing tool on their agent (rename it, fix its description, change a webhook URL/headers/auth, add or edit a request-body parameter or response assignment, adjust the timeout, or flip a client tool's expects_response / execution mode), or when a tool edit won't save or throws schema_mismatch / validation / not_found.
---
Editing a tool fails a lot, almost always because the patch violated a per-type schema rule or dropped a field the caller never read. The fix: read the current tool first, send a minimal but complete-per-object patch. All calls use `xi-api-key: $API_KEY` against `https://api.elevenlabs.io`.
## 1. Read the current tool first
Do not edit blind; the failures come from patching a shape you never read.
- `GET /v1/convai/tools` to get the exact tool id and its TYPE (webhook / client / code). The type determines the config shape.
- `GET /v1/convai/tools/{tool_id}` to read the tool's CURRENT saved config: name, description, `request_body_schema` properties, `request_headers`, auth, response `assignments`, `response_timeout_secs`, and (for client tools) `expects_response` / `execution_mode`.
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` to confirm the tool is attached (`tool_ids` / per-node `additional_tool_ids`) and to see which nodes reference it before you rename anything.
## 2. Update with the full config
```
PATCH /v1/convai/tools/{tool_id}
```
`PATCH` REPLACES the config, so send the config for the type you read in step 1 with your changes applied. Change only what the user asked, but include the FULL object for any nested field you touch: if you edit one property in `request_body_schema`, resend the whole `request_body_schema`, not just the one property. Partial nested objects are the usual source of silent drops and validation errors.
A SYSTEM tool (`transfer_to_number`, `end_call`, `language_detection`, `skip_turn`, `update_state`, etc.) is NOT edited this way; its config lives in the agent config, so edit it via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` targeting the relevant path.
## 3. Schema gotchas
Webhook / code tools (`request_body_schema`):
- Each property needs EXACTLY ONE of: a non-empty `description`, `constant_value`, `dynamic_variable`, or `is_system_provided`. Two (or zero) is a validation error. When editing, do not leave the old `dynamic_variable` while adding a `description`.
- Array properties REQUIRE an `items` schema and MUST NOT carry `constant_value`.
- URLs must be STATIC, no `{{variables}}` in the path; move dynamic values into `request_body_schema`.
- `request_headers` is an OBJECT, not an array. For secrets use `{"secret_id": "..."}` (list via `GET /v1/convai/secrets`); do not paste a raw key when the original used a secret id.
- Response `assignments` use dot-notation `value_path` (`data.items.0.id`, NOT `data.items[0].id`; bracket indexing silently returns null).
- `response_timeout_secs`: default 20, max 120.
Client tools:
- `expects_response` (bool) and `execution_mode` (`immediate` / `post_tool_speech` / `async`) are the two fields people edit and get wrong. If `execution_mode` is `async`, the tool is fire-and-forget and `expects_response` must be false.
- Parameters follow the same one-of-four value-source rule.
All types:
- Renaming: the tool `name` is referenced by the LLM and by any node prompts or edge conditions that call it by name. After a rename, prompts still using the old name will not fire the tool. Flag this and offer to update the referencing prompts in the same pass via the agent config.
- Do not invent new schema keys.
## 4. Recovery
- `schema_mismatch`: the patch shape does not match. Re-`GET` the tool, diff against what you sent, and resend the full nested object for the field you touched. Do not retry the identical body.
- `validation`: a per-field rule above was broken (two-of-four on a property, array with `constant_value`, async + `expects_response`, timeout > 120, `{{var}}` in URL). Fix that field and resend.
- `not_found`: a stale or wrong tool id, or the tool is not on this branch. Re-`GET /v1/convai/tools` for the current agent/branch and use the exact id.
## 5. Follow-ups
If you renamed the tool or changed its parameters, update referencing node prompts or edge conditions via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`, and re-check any tool test that targeted the old schema (a sim test edit is delete + recreate; see the create-tool-test skill). If the underlying problem is that the tool is not being CALLED or returns no result at runtime (not a config-edit failure), use the troubleshoot-tool-errors skill instead.
architect-edit-string-fields5.37 KB
---
name: architect-edit-string-fields
description: Use when changing wording inside a long text field on an agent (system prompt, first message / greeting, tool description, workflow node prompt, or procedure content) - "add this to the prompt", "change this line", "append a rule", "reword the greeting", "fix the tool description" - or when a previous edit clobbered or failed to save the field.
---
# Edit large string fields
Change one long text field by reading it, editing the whole string locally, and writing it back. Do not resend a big multi-field config blob to change wording. Broad config PATCHes are the high-failure path, and resending a big blob risks silently dropping the rest of the field.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID`.
The general move is read, modify, write: GET the field, produce the full new string (append a line, replace a passage, or rewrite wholesale), then PATCH the whole field back. There is no browser find/replace and no read-before-edit gate here. You always have the current text because you just fetched it, so compute the exact new string yourself and never edit from memory.
## System prompt and first message
Read the field from the branch-scoped agent:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" -H "xi-api-key: $API_KEY" \
| jq -r '.conversation_config.agent.prompt.prompt' > prompt.txt
```
Edit `prompt.txt` to the full intended text (append your rule, replace the passage, or rewrite it), then PATCH the whole string back. A partial-merge PATCH commits to branch HEAD and returns a new `version_id`:
```bash
curl -s -X PATCH "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID&version_description=prompt-edit" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d "{\"conversation_config\":{\"agent\":{\"prompt\":{\"prompt\":$(jq -Rs . < prompt.txt)}}}}"
```
`jq -Rs .` JSON-encodes the whole file, so quotes, arrows, dashes, and newlines are handled for you. For the first message, read and write `conversation_config.agent.first_message` the same way.
## Tool description
The description lives on the tool, not the agent. Get the tool, edit its `description`, and PATCH the full tool config back (the PATCH replaces the tool):
```bash
curl -s "https://api.elevenlabs.io/v1/convai/tools/$TOOL_ID" -H "xi-api-key: $API_KEY" > tool.json
# edit the description field in tool.json to the full new string
curl -s -X PATCH "https://api.elevenlabs.io/v1/convai/tools/$TOOL_ID" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
--data-binary @tool.json
```
## Workflow node prompt
The node prompt lives in the workflow inside the agent config. Only `override_agent` nodes and prompt-type `say` nodes have an editable prompt. Read `workflow` from the branch-scoped agent, find the node by id, set its prompt string to the full new text, and PATCH the workflow back through the config. If you cannot find an editable prompt on the node, you have the wrong node id or a non-editable node type.
## Procedure content
Procedure content does not go through the agent-config PATCH. It goes through the procedure draft, then a publish. This is two steps, and there is no direct PATCH on a procedure (that returns 405).
```bash
# 1. Read the current procedure body.
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID" \
-H "xi-api-key: $API_KEY"
# 2. Stage the full new content as a draft (send name, content, type, trigger).
curl -s -X PATCH \
"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID/draft" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"name":"...","content":"...","type":"free_form","trigger":"..."}'
# 3. Publish pending drafts by committing the agent with any partial-merge body.
# Re-sending the branch's current prompt unchanged is the simplest no-op publish.
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" -H "xi-api-key: $API_KEY" \
| jq -r '.conversation_config.agent.prompt.prompt' > current_prompt.txt
curl -s -X PATCH "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID&version_description=publish-procedure" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d "{\"conversation_config\":{\"agent\":{\"prompt\":{\"prompt\":$(jq -Rs . < current_prompt.txt)}}}}"
```
Fetch the current prompt scoped to `$BRANCH_ID` (not main) before the no-op publish so you do not clobber other branch-local differences. A full-content draft PATCH replaces the whole body, so it is cleaner than trying to splice a passage in place. The draft PATCH cannot change a procedure's `type`.
## Verify
After any write, re-read the same field and confirm the new text is present and nothing else was disturbed. For procedures, confirm the `version_id` changed after the publish step. These edits land on the current branch and ship only when the branch is merged, so remind the engineer to review and test the new behavior before merging.
## When it is not a string-field edit
A change to language, TTS model, LLM, or any non-text single field is not a string edit. Route it through the config PATCH flow and remember the language and TTS-model cross-field constraint. See the `architect-update-config-safely` skill.
architect-edit-workflows4.6 KB
---
name: architect-edit-workflows
description: Use when the user wants to build or change an agent's workflow or a specific node ("add a node", "connect these nodes", "route to a phone transfer", "add a condition/edge", "delete this node", "why does deploy say duplicate edge"), or a workflow change is failing validation.
---
An agent's workflow (its nodes and edges) lives inside the agent config. Over REST you read it from `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` (the `workflow` object) and write whole-config changes with `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. There is NO granular per-node/per-edge endpoint from the API. If the user needs interactive node-by-node editing on a canvas, that is only in the web app; say so plainly and either make the whole-config edit here or point them there. All calls use `xi-api-key: $API_KEY`.
Workflow writes are error-prone. Go slow, change one thing at a time, and read the current graph before every write.
## 1. Read the current graph first, always
`GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` and read the `workflow` object. Nodes and edges are keyed maps (not arrays): every node has a `type`, a `position`, and an `edge_order`; the graph always contains `start_node`. You need existing node ids to connect, update, or delete, and you need to see which pairs already have an edge. Also `GET /v1/convai/tools` for the real tool ids: expression edges and tool nodes reference tool ids, and a guessed id fails validation.
## 2. Make one change at a time
Take the workflow you read, apply a single change, and PATCH the whole `workflow` back. Do not hand-build a partial workflow from memory and do not drop nodes, edges, or nested config you did not mean to change: omissions are treated as deletions and cause `schema_mismatch`. After each successful write, re-read the config to confirm the change landed before the next one.
## 3. Cross-node constraints (each is a common failure)
1. ONE edge per node pair. Two edges between the same pair cause a "Duplicate edge found" error on deploy. For a bidirectional transition, put `forward_condition` (A to B) and `backward_condition` (B to A) on the SAME single edge. Never create a second edge for the return direction. This is the most common failure.
2. Terminal nodes have NO outgoing edges. `end` and `phone_number` (phone transfer) nodes are terminal; an edge out of one, or a `forward_condition` from one, fails validation. A phone_number transfer that fails just ends the call; if the user wants retry or fallback, recommend the `transfer_to_number` system tool instead.
3. Expression edges vs LLM conditions. Use an expression edge for deterministic routing off a tool result or dynamic variable (route on whether a lookup returned a match, or on `{{account_type}}`). Use an LLM/prompt condition only for genuine judgment. Vague LLM conditions cause wrong-node routing; write tight conditions that reference concrete signals ("lookup returned a matching record", not "customer is identified").
4. Conditions must be mutually exclusive. Overlapping conditions on a node's outgoing edges make the model fire the wrong one. Keep each outgoing condition disjoint.
5. Node ids and positions are required and consistent. A connection needs source and target ids that already exist in the current graph (why step 1 is mandatory). A new node needs a `position` ({x, y}); omitting it or reusing an existing id causes validation errors.
6. A transfer (`standalone_agent`) node keeps `agent_id` and `node_id` as FLAT fields, siblings of `type`/`position`; there is no `transfer_destination` wrapper. Set `agent_id` (and leave `node_id` null) to transfer to a different agent; set `node_id` (and leave `agent_id` null) to transfer within the same agent. The "a target must be selected" error fires only when both are empty. The node also needs `delay_ms` (int >= 0), `transfer_message` (string or null), and `enable_transferred_agent_first_message` (bool).
7. On an agent with `override_agent` nodes the effective prompt lives IN the nodes, so editing the base agent prompt often has no effect on the conversation or 422s. Change the node's prompt in the workflow, and tell the user which node owns the text so they are not surprised a base-prompt edit changed nothing.
8. Tool scoping is separate from the graph. If a node needs a tool, set `tool_ids: []` at the base agent and grant tools per node via that node's `additional_tool_ids`, not by adding edges.
## 4. After editing
Re-read the config to confirm the change landed, and remind the user to test each transition path end to end (a simulation test is the reliable way; see the create-simulation-test skill).
architect-explain-test-runs4.39 KB
---
name: architect-explain-test-runs
description: Use when the user asks why a test or test run passed or failed, why one run in a repeat set differs, which procedures or tools were used during a run, or wants a summary of test suite history. Fires on "why did 1 of the 3 runs fail", "why is this test failing", "which procedures were used in that run", "explain this simulation test result", or "summarize my test suite history".
---
# Explain why a test run passed or failed
The user wants for a test run what they already get for a conversation: which procedures fired, which tools were called, and what actually happened turn by turn. Read the test invocation, find the divergence, and answer with it.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY` and `$AGENT_ID`.
## Order of investigation
1. **List recent runs.** Fetch recent suite invocations, optionally filtered by agent, to find the invocation in question and to compare runs across time. To surface "frequent offenders", compare recent invocations and identify which specific tests fail repeatedly; those need fixing most.
2. **Read the invocation.** `GET /v1/convai/test-invocations/{suite_id}` returns `test_runs[]`, each with `status`, `condition_result.result`, and `condition_result.rationale`. The rationale per run is often specific enough on its own. For most "why did X fail" questions this is enough; stop here if the rationale already names the divergence.
3. **Read the full transcript only when the rationale is too vague.** The invocation also carries the turn-by-turn transcript per run: messages, tool calls (including the procedure-start calls, so you can see which procedures fired), and tool errors. This is what answers "why did it fail" or "which procedures ran in each pass and fail run". On a suite with many tests and a `repeat_count` above 1, this payload can be large enough to overflow context. Check the test and run count first; if it covers more than a handful of tests or repeats, do not dump the full detail. Instead tell the user the suite is too large and offer to (a) look at just the failing test's rationale, or (b) point them to the Tests tab in the UI.
4. **Read the test definition** for the criteria, scenario, and mocking config when you need to see exactly what was asserted.
5. **For a repeat set, compare a passing run against a failing run of the same test.** The divergence is the answer: if the runs called different procedures, that is a routing problem; if they called the same procedures and tools, that is an eval or non-determinism problem.
## Answer with the divergence, not the criterion
"It failed because it did not meet criterion X" restates what the user can already see and answers nothing. Instead:
- Name the specific turn where the failing run differs and what it did differently. "In the failing run the agent updated the ticket to solved before the customer-facing reply; in the passing runs it replied first."
- Quote the clause of the criterion that was violated, not the criterion's name.
- Lead with the answer in the first sentence and keep the whole explanation under about six lines of prose.
## When the config is identical
Same config and input, different outcome, is the most common version of this question. Say so directly: the model made a different choice on that run, which is LLM non-determinism, not a config bug. Then give the actionable fix: tighten the criterion if it is ambiguous at that boundary, or move the requirement into a deterministic procedure if the ordering must be guaranteed. Never present a sampling difference as though the failing run had a distinct root cause.
## Summarizing suite history
When the user wants a history summary, present the overall pass/fail counts, a dedicated section for critical issues (critical tests are often tagged in the title, otherwise infer from context and content), and a section for recurring failures (tests that failed across multiple recent runs). Then diagnose specific failures from the invocation detail and suggest concrete fixes.
## Honesty
If a read errors, a transcript is truncated (check for a truncation flag and the total turn count), or a suite is too large to fetch in full, say which read failed or what was cut off, and offer to point the user to the Tests tab in the UI. Never fabricate a turn-by-turn account you could not read, and never present the evaluator's rationale as though it were the transcript.
architect-explore-agent2.76 KB
---
name: architect-explore-agent
description: Use when starting a task that needs an agent's current setup, or when another skill tells you to explore the agent first, before making recommendations or changes to a ConvAI agent.
---
# Explore an agent's context
Before recommending or changing anything, gather just enough of the agent's current configuration to act. Fetch narrowly. Do not pull the whole config every time.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID` where relevant.
## 1. Read only what the task needs
Get the agent scoped to the branch:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" \
-H "xi-api-key: $API_KEY"
```
This one response carries the whole config: `name`, `conversation_config.agent.prompt.prompt` (system prompt), `conversation_config.agent.prompt.llm`, `conversation_config.agent.language`, `conversation_config.tts.model_id` and `voice_id`, `conversation_config.agent.prompt.tool_ids`, guardrails, and the workflow. Read the fields the task needs and ignore the rest. The whole config is large, so do not re-fetch it to read one leaf you already have.
## 2. Add narrow follow-up reads only when the task touches them
- To inspect a specific tool, take its id from `conversation_config.agent.prompt.tool_ids` and get just that tool: `GET /v1/convai/tools/{tool_id}`. Do not re-fetch the agent to read one tool.
- For procedures, list names and triggers with `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures`, then get one body by id with `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/{procedure_id}`.
- For past conversations, `GET /v1/convai/conversations?agent_id=$AGENT_ID`, then `GET /v1/convai/conversations/{conversation_id}` for a transcript. These carry customer PII and analysis. Do not copy them into other systems or logs, and respect zero-retention-mode accounts. For anything beyond a glance at one or two calls — pagination, which fields the list payload actually carries, aggregating a window, sampling — use `architect-review-live-calls`, which owns that pattern.
Make independent reads in parallel where they do not depend on each other, each still narrowly scoped.
## 3. Identify the active setup
From what you fetched, note the active LLM, voice, TTS model, language, and any configured tools or custom guardrails.
## 4. Summarize, then ask
State the current setup in one concise sentence, then ask what the engineer wants to improve before pulling anything more. An empty workflow (only the start node with no edges) is the default state, not a custom configuration. If the agent has procedures but no workflow nodes, describe it as a procedure-driven agent with no workflow configured.
architect-generate-agent2.54 KB
---
name: architect-generate-agent
description: Use when the user asks to generate, build, or create a new agent from a description ("create an agent for X", "build a chatbot that does Y"). There is no one-shot generate endpoint over REST; ask clarifying questions, then build via create and refine via patch.
---
# Generate a new agent (there is no one-shot generate endpoint)
The web Architect has a `generate_agent` action. Over the ConvAI REST API there is NO one-shot generate endpoint. Do not pretend there is. Instead, gather requirements, create a starter agent with `POST /v1/convai/agents/create`, then refine it with `PATCH` (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`).
## Step 1: pause and ask clarifying questions
Before creating anything, understand what the user actually needs. Ask targeted questions to gather:
- Use case and context: what specific problem does this agent solve? Who are the end users?
- Scope and capabilities: what should the agent do, and what should it NOT do?
- Tone and personality: how should it sound (formal, friendly, technical)?
- Integration needs: does it need to call external APIs, access a knowledge base, or trigger workflows?
- Success criteria: how will the user know it is working well?
- Constraints: any compliance, language, or channel requirements?
## Step 2: synthesize and confirm
Summarize your understanding back to the user in 2-3 sentences and ask "Is this what you're looking for?" This prevents building the wrong agent.
## Step 3: create the agent
Once aligned, create the agent with `POST /v1/convai/agents/create`. Set the name and an initial `conversation_config` built from the synthesized requirements: a system prompt capturing the use case, scope, tone, and constraints, a suitable first message, and an LLM chosen for the task (see the `architect-llm-selection` skill for region/compliance/latency tradeoffs). The response returns the new `agent_id`.
## Step 4: refine
Iterate with `PATCH /v1/convai/agents/{agent_id}?branch_id={b}` partial-merge bodies to tighten the prompt, add tools, or adjust config. For step-by-step behavior, add procedures (see the `architect-manage-procedures` skill). For knowledge, attach knowledge base documents via `POST /v1/convai/knowledge-base/text` or `.../url`. For post-call extraction and grading, see the `architect-post-call-data` skill.
## Step 5: review and next steps
Tell the user the agent is created and offer next steps: review the system prompt, set up procedures, test with a simulation, secure it for production, or customize the workflow.
architect-get-agent-ci-green15.3 KB
---
name: architect-get-agent-ci-green
description: Use when an agent's test suite is red and needs to be driven to green. Fires on "get CI green", "the test suite is failing", "run the tests and fix them", "triage these test failures", "why is my agent suite red", or after attaching tests when many runs fail at once. Covers running a suite, polling it, classifying failures by cause, and fixing the ones that are test bugs rather than agent bugs.
---
# Get an agent's test suite green
A red suite is usually not one bug. It is several unrelated causes stacked together, and most of them are broken *tests*, not a broken agent. Triage by cause first; fixing blind wastes whole runs, since each simulation takes minutes.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, `$BRANCH_ID`.
## 1. Establish a baseline before changing anything
You cannot tell a fix from noise without a starting number.
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" \
-H "xi-api-key: $API_KEY" > agent.json
```
The suite is at `platform_settings.testing.attached_tests` — a mix of individual `test_` ids and `tfld_` folder ids. **Only these are this agent's CI.**
That distinction matters more than it looks. On a real Architect suite, 56 tests ran but only 3 were attached to the agent; the other 53 came from inherited folders and belonged to entirely different agents (a Zendesk support agent, a delivery agent, a sales agent). They failed because they asserted another agent's behavior. No amount of fixing makes them pass here, and "fixing" them means editing another team's tests.
So before triaging, split failures into:
- tests attached to **this** agent — your problem
- tests that arrived via a shared folder and target another agent — not your problem; report them and leave them alone
Fetch each test body with `GET /v1/convai/agent-testing/{test_id}`. A 404 on an attached id means a stale attachment pointing at a deleted test — worth reporting, and it cannot pass.
## 2. Run the suite and poll
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"tests": [{"test_id": "test_abc"}], "branch_id": "'"$BRANCH_ID"'", "repeat_count": 1}'
```
The response `id` is the suite id. Poll it, counting statuses:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/test-invocations/$SUITE_ID" -H "xi-api-key: $API_KEY"
```
Poll in the **background** on a 20-30s interval, not a foreground sleep. Simulations take minutes and a large suite takes far longer; a foreground wait blocks the session and times out. Stop when no `test_runs[].status` is `pending`.
To find a run started from the UI, list recent invocations (note this is a top-level endpoint; the `/agents/{id}/test-invocations` form does not exist):
```bash
curl -s "https://api.elevenlabs.io/v1/convai/test-invocations?agent_id=$AGENT_ID&page_size=10" \
-H "xi-api-key: $API_KEY"
```
## 3. Classify every failure before fixing any
Read `condition_result.rationale.messages` plus the transcript at `agent_responses[]`. Sort each failure into one of four buckets — they have different fixes and very different costs.
**A. Unmocked tool call.** Transcript shows `Error: no mock matched for tool 'X'`. A test bug, and normally the largest bucket. In one real triage this was 32 of the failures, all from a single tool. Fix per the `architect-mock-all-tools` skill: mocks key on **tool ID**, and the ID must be one on this agent. Name keys and stale IDs are silently ignored.
**B. 60s timeout.** `Timed out after 60s waiting for the agent to produce its next turn`, with an empty transcript. This is a hang, not a verdict — the agent never got started. For DOM/Architect-style agents the usual cause is an empty `chat_history`. Seed one agent turn and raise `simulation_max_turns` to at least 6.
If seeding does not fix it, the timeout can be **content-driven** rather than config-driven. Verified by A/B: two tests with byte-identical mock config, seed, and turn count behaved differently purely by scenario text — "Which model is this agent on?" passed in one turn, while "Add a client tool ..." timed out with zero tool calls on every attempt. Heavy authoring requests (creating tools, generating agents) can exceed the 60s per-turn budget before the agent emits anything. Swap in a trivial scenario on the same config to tell the two apart: if the trivial one passes, the config is fine and the request itself is too slow — a platform limit to report, not a test to keep retrying.
**C. Wrong-agent test.** Success conditions reference tools or flows this agent does not have. Unfixable here by design. Report it, do not edit it.
**D. Genuine agent misbehavior.** Mocks all matched, the agent ran the flow, and it still did the wrong thing. **This is the only bucket that should change the agent.** Everything else is test repair.
**E. Stale architecture.** A `tool`-type test asserts a tool that the agent no longer routes through, so every run reports `No tool with name 'X' was called`. The agent is correct and the assertion is obsolete.
This is easy to mistake for a mass regression because it fails loudly and in bulk. In one triage, 46 of 108 failures were a single instance of this: every test asserted `execute_agents_tool`, but a migration had promoted the catalog tools to direct ConvAI tools, so that dispatcher is no longer called. Nothing was broken; the tests encoded the previous architecture.
The signature is a large group of same-type tests failing with an identical "tool not called" message and a shared `referenced_tool` id. Check whether that tool is still on the agent's `tool_ids` and whether a recent migration changed how it dispatches. The fix is to rewrite the assertions against the tool now called, or retire the tests — a product decision, so surface it rather than silently deleting.
Rewriting a meta-tool assertion to a direct one is mechanical. The old envelope asserted `tool_name` plus `args.*` paths; the direct call drops the `tool_name` row, strips the `args.` prefix from every remaining path, and repoints `referenced_tool.id` at the real tool. Watch for two follow-on traps: a migration that *splits* one tool by type (`create_tool` → `create_client_tool` / `create_webhook_tool`, `create_test` → `create_llm_test` / `create_simulation_test` / `create_tool_test`), where the right successor is inferable from the asserted arg paths; and stale `execute_agents_tool` calls sitting in the test's own seeded `chat_history`, which teach the agent the obsolete pattern in-context and must be rewritten too.
**Before rewriting, check the test type can even observe the call.** A `tool`-type test grades a single agent turn — exactly one tool call. If the agent opens with `start_procedure` (or any preamble) and calls the target on a later turn, the test can never pass no matter how correct the assertion is. Setting `check_any_tool_matches: true` widens matching from "the first call" to "any call in the turn", which is worth doing because it converts a misleading `Expected tool 'X' but agent called 'start_procedure'` into an honest `No tool with name 'X' was called` — but it does not add turns. In one suite, all 50 tool tests produced exactly one call and 37 opened with `start_procedure`, so they were structurally unpassable as `tool` tests.
### Converting a tool test to a simulation test
For a procedure-first agent this is the real fix, and it is mechanical. Build the simulation from the tool test's own fields:
- `simulation_scenario` — replay the tool test's seeded user turns as second-person instructions ("Say exactly: '<first user message>'. If the agent asks for confirmation, reply '<second user message>'. Then say you are done.").
- `success_conditions` — one condition per assertion: the `referenced_tool` name becomes "The agent calls the X tool", and each `parameters[]` entry becomes a sentence about its `path` and expected value.
- `chat_history` — a single seeded agent greeting; `simulation_max_turns` around 14.
- Mocks — the full ID-keyed superset, as always.
The payoff is real: a converted test ran 11 tool calls and reached the target the single-turn version could never see.
Two things to get right, both learned the hard way:
- **Assert intent, not an exact string, where the agent has legitimate freedom.** One converted test set the LLM to `gemini-2.5-flash` on one run and `gemini-2.5-flash-preview-05-20` on the next; both are valid ids for "Gemini 2.5 Flash", so an `exact` match is flaky by construction. Write "a gemini-2.5-flash* variant" instead. Same for tools split by type — accept `update_agent_config` or `patch_agent_config` when either is correct.
- **Re-run a converted test at least twice** before trusting it. Single-run greens hide exactly this nondeterminism.
### Assert the outcome, not the tool
The largest source of false failures is a criterion naming one specific tool when the agent legitimately reached the same outcome another way. Expect this to outnumber real bugs.
Three forms, all fixed by rewriting the criterion rather than "fixing" the agent:
- **A purpose-built tool beats the generic one.** Where a catalog offers both a specialized writer and a general config writer, the agent will usually pick the specialized one. A criterion demanding the generic writer then fails a correct run. Name the acceptable set ("any tool that writes X: A, B or C") and assert the resulting state instead of the call.
- **Assertions on fields the schema does not require.** When a tool is split by type, the identity moves into the tool name and any old type argument becomes redundant, so the agent omits it. Criteria asserting such a field fail against correct calls. Read the live schema (`GET /v1/convai/tools/{id}`, check `parameters.required`) and delete assertions on anything optional.
- **Values differing in case or form.** Decide whether exact casing is genuinely part of the contract; if not, say "(case-insensitive)" in the condition rather than failing a correct call.
The discipline: when a run fails, read what the agent actually did **before** editing anything. If it achieved the user's goal by a reasonable route, the criterion is wrong. Only conclude "agent bug" once the criterion describes an outcome the agent genuinely failed to reach.
### The scenario must justify the assertion
Conversion carries old assertions forward verbatim, and some were never justified by the prompt. If the scenario asks for an everyday outcome but the criterion demands a specific internal structure, the agent will satisfy the request the obvious way and fail. A test whose prompt does not imply its assertion is a broken test, not a failing agent.
Read the scenario and criteria together. If a reasonable user reading that prompt would not expect the asserted outcome, either make the prompt ask for it plainly or relax the criterion to what the prompt actually implies. The same applies to the kind of artifact you expect: if the scenario does not say which type to create, the agent may reasonably create a different one.
Only after A, B, C, and E are cleared can you trust bucket D. Fixing the agent to satisfy a test that was failing for reason A or E actively makes the agent worse.
## 4. Watch for implausible mock data masquerading as an agent bug
A failure looks like bucket D — mocks matched, agent declined to act — but the real cause is mock *content*. Agents reason about values, so a well-formed mock with an unrealistic value gets rejected.
Seen in practice: a ticket mocked as `TRIAGE-4821` when the platform's ids are `agtqa_`-prefixed. The agent judged it malformed, refused the lookup, and every success condition failed. Changing only the id to `agtqa_3901...` turned the same test green. Before concluding the agent is wrong, check that mock values match production shape.
## 4b. Knowledge failures: verify the fact before changing the agent
When a test fails because the agent stated a wrong number or denied a real feature, do not patch the prompt from the test's expectation. The test may be the stale one. Find the source of truth in the codebase first — the enforced constant, with a file:line — and only then decide which side is wrong. Parallel targeted subagent searches are well suited to this: each fact is an independent lookup.
Three outcomes, each with a different fix:
- **Test right, agent wrong** — add the fact to the prompt, using the verified value.
- **Test wrong, agent right** — fix the test. Product facts drift, and a test written against an old figure will keep failing a correct agent.
- **Sources genuinely conflict** — surface it rather than picking a side. When config, docs, and the test disagree, that needs a product ruling, not a prompt edit.
Watch for a single systemic cause behind many "wrong number" failures. Limits often differ between surfaces — the web app, the public API, and per-plan tiers — and an agent that reaches for the wrong one will miss a whole cluster of facts at once. Fixing that framing once resolves many individually-reported failures, so before writing six separate fact patches, check whether they share one root cause.
Read the judge's rationale closely too: it often quotes the knowledge base directly and tells you which side is stale.
## 5. Fix in place, then re-run only what you changed
Repair a simulation test with `PUT /v1/convai/agent-testing/{test_id}`. It edits in place and keeps the id, so attachments and folder placement survive. Send the complete object — omitted fields are not preserved — and note that `PATCH` returns 405.
Do not delete and recreate to apply a fix. A new test gets a new id, which silently drops it from the agent's suite, so the next full run covers less than you think while looking greener. A 404 from `PUT` means a wrong id, not a missing capability.
Bulk repair works well because the two mechanical buckets share one payload shape: seed `chat_history`, raise `simulation_max_turns`, and attach the full ID-keyed mock set. Applying that to every A and B failure at once is a single pass — 38 tests updated in one batch in a real triage, all returning 200.
Deleting a test you do not own is destructive; confirm with the owner first, and never delete another team's tests to make your number go up.
Re-run just the changed tests while iterating. Save the full-suite run for final verification.
## 6. Verify green means green
Two failure modes hide behind a passing suite. Check both.
**A test can pass with every mock broken.** One verified run passed on the agent's own reasoning while all seven of its mocks returned `no mock matched`. Grep every transcript for `no mock matched` even when the suite is fully green; any hit is a latent failure that will surface the moment the agent's path changes.
**Passing once is not passing reliably.** Agents take nondeterministic paths, so a run can pass while leaving tools unmocked that a later run will hit. Observed directly: a repaired test passed, then on re-run the agent reached `patch_agent_config` and `update_agent_config`, which nothing had mocked. Mock the superset of tools the agent *could* call, not just those it happened to call once, and re-run a repaired suite at least twice before declaring it green.
Report the outcome as a delta with both numbers (`1/6 → 5/5`), name what stayed red and why, and state plainly which failures were test bugs versus agent bugs.
architect-llm-selection5.7 KB
---
name: architect-llm-selection
description: Use when choosing or recommending an LLM for an ElevenLabs agent, or answering "which model should this agent use", questions about region-restricted models, HIPAA/PCI/ZRM model eligibility, or latency/intelligence/cost tradeoffs between models.
---
# LLM selection for agents
Resolve constraints in this sequence. Each step narrows the candidate set; never widen it later. Everything is over the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`).
1. Explore the agent. Read the current LLM, workflow, and tools with `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` first.
2. Region (infrastructural, non-negotiable). Determine deployment topology. EU and India customers hit isolated deployments with restricted model lists that apply to all tiers. A model hidden by region cannot be enabled; do not offer to.
3. Compliance entitlements (contract-gated). If the workspace requires HIPAA, PCI, or ZRM, filter to compliant models only. These are workspace-level flags managed by sales/CSMs, never self-serve.
4. Capability fit (intelligence / latency / cost). Within whatever survives steps 1-3, recommend on task requirements.
If a constraint in step 2 or 3 eliminates the model the user wanted, say so plainly and route to their CSM or a support form. Do not present a capability recommendation that violates a hard constraint.
## Step 1: region gate
| Region | Rule |
| --- | --- |
| US / default | Full model catalog, subject to compliance gates. |
| EU-isolated | Restricted list. Preview Gemini models and several others are hidden. |
| India-isolated | Heavily restricted: only `custom-llm`, `glm-45-air-fp8`, `gpt-4o`, `qwen3-30b-a3b`, `qwen36-35b-a3b`, `speech-engine`. |
Restriction is by deployment topology, not a toggle. Never offer to enable a region-hidden model.
## Step 2: compliance gate
Only relevant if the workspace has the corresponding entitlement. You cannot see workspace flags, so never assert an entitlement is active; the safe line is "your CSM can confirm this for your workspace." Never offer to enable HIPAA, PCI, or ZRM; they require contract changes.
- HIPAA requires all three: enterprise tier with `force_logging_disabled=true`, an LLM in the HIPAA-compliant set, and a signed BAA (a contract artifact you cannot verify; route BAA questions to sales). Approved families include Claude 3.5/3.7 Sonnet, Claude 3 Haiku, Claude Haiku 4.5, Claude Opus 4.7, Claude Sonnet 4 / 4.5 / 4.6, `custom-llm`, Gemini 1.5/2.0/2.5 flash and pro variants, `gemini-3.1-flash-lite`, and `speech-engine`. Models outside the list are not HIPAA-eligible regardless of tier; preview/GA-pending Gemini models are excluded.
- PCI is deny-by-default when on (`pci_compliance_required`). Only pre-approved telephony providers, webhook domains, MCP URLs, and integration IDs work. Adding an integration requires CSM review; offer a support form, do not promise self-serve.
- ZRM disables audio/transcript storage; effective only when `force_logging_disabled=true`. Custom LLM plus ZRM additionally needs `is_convai_custom_llm_with_zrm_allowed`. No Stripe SKU; part of the enterprise contract.
## Step 3: capability fit
Recommend within the surviving candidate set. Weigh three axes against the agent's job.
- Latency matters most for voice. In an STT to LLM to TTS loop, LLM time-to-first-token sits on the critical path and is the dominant lever on perceived delay. For real-time agents, bias toward the fastest tier that clears the quality bar.
- Lowest latency: `gemini-3.1-flash-lite` / `nano` / small `qwen` / `*-mini` tiers and `speech-engine`.
- Balanced: `gemini-3.5-flash`, `gpt-4o`, `gpt-5-mini`, `claude-haiku-4-5`.
- Highest capability, higher latency: `claude-sonnet-4-6`, `claude-opus-4-7`, `gpt-5.x`, `gemini-3.x-pro` (where region/compliance permits).
- Intelligence. Reserve the top tier for genuinely hard reasoning, complex multi-tool orchestration, or nuanced instruction-following. Most transactional voice agents (booking, triage, FAQ, routing) run well on a mid tier and feel snappier. Over-provisioning intelligence usually costs latency the user will notice for quality they will not.
- Cost. Per-token price scales steeply with capability tier. Match tier to task; do not put a frontier model behind a deterministic IVR flow. For high-concurrency deployments, latency and cost compound.
### Quick heuristic
| Agent type | Priority | Typical pick (region/compliance permitting) |
| --- | --- | --- |
| Real-time phone / high concurrency | Latency, cost | `gemini-3.1-flash-lite` (ultra-low latency) or `gemini-3.5-flash` |
| Complex reasoning / multi-tool | Intelligence | `claude-sonnet-4-6`, `gpt-5.x`, `claude-opus-4-7` |
| HIPAA voice agent | Compliance, then latency | `claude-haiku-4-5`, `gemini-3.5-flash` (both in approved list) |
| India deployment | Region first | `gpt-4o` or `qwen36-35b-a3b` (catalog is tiny) |
| Custom/self-hosted model | Control | `custom-llm` (check ZRM + region eligibility) |
## Apply the choice
Set the model with `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` and body `{"conversation_config":{"agent":{"llm":"<model-id>"}}}`. This is branch-scoped and ships on merge.
## Output discipline
- State the binding constraint explicitly ("In the EU deployment, X is not available, so among allowed models...").
- Give one primary recommendation plus a fallback, with the reason (latency vs intelligence vs cost).
- For anything contract-gated, end with the route: CSM for entitlement confirmation, a support form for new integrations/documents, and https://compliance.elevenlabs.io/ for attestations and signed DPA/BAA/sub-processor requests.
- Never claim an entitlement is active, never offer to enable a contract feature or a region-restricted model.
architect-manage-procedures9.96 KB
---
name: architect-manage-procedures
description: Use when you want to add, edit, rename, inspect, or list an agent's procedures, or when a procedure create/edit call fails with schema_mismatch, validation, not_found, a 405, or a 422 missing `type`. Also use when a newly created procedure comes back empty or 404s on read. Covers the list/get/create/edit-draft/publish flow over the ConvAI REST API.
---
# Manage agent procedures reliably (list / get / create / edit / publish)
Procedures are their own data, not part of the agent config. There is no agent-config path for them and no way to reach them by patching `conversation_config`. Everything here is branch-scoped and goes through the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`). You need `$AGENT_ID` and the branch id `$BRANCH_ID` you are working on.
For the craft of writing a good trigger and clear steps, see the `architect-structured-procedures` skill and the procedure-authoring guidance. This skill is about calling the endpoints correctly.
## 1. Read current state first
Before creating or editing anything, read what exists:
- List every procedure on the branch: `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures`. Each entry carries `procedure_id`, `version_id`, `name`, `type` (`free_form` or `deterministic`), `trigger`, and `has_draft`. It does NOT include `content` — the list is an index, not a way to read bodies. Do not measure a procedure's content length from this response; an absent `content` field will read as empty for every procedure and look like data loss when nothing is wrong. You cannot edit a procedure without its `procedure_id` from here.
- Read one procedure's body with `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID`. This is the only way to get `content`, and you need it before an edit so you do not blow away the rest. If a draft exists for that procedure, this GET returns the DRAFT content, not the published content.
- This GET returns 404 for a procedure whose content has never been published, even when the id is correct and the procedure appears in the list. That is the expected state right after a create (see section 2), not a wrong id. Use the list to confirm the procedure exists and check `has_draft`.
- Confirm the branch you are on with `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`, since procedures are per-branch.
If the user asked to edit "the X procedure", match X against the `name` values from the list and confirm the exact `procedure_id` before writing. Do not assume an id.
## 2. Create vs edit: the draft then publish model
Writes stage a DRAFT on the branch. A draft is not live until a separate publish step commits it into a new version.
- CREATE a new procedure: `POST /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures` with `name` and `content`. `type` defaults to `free_form`; pass `type: "deterministic"` only when writing the JSON flow form.
**Creating is two calls, not one.** The POST registers the procedure and stores the `name`, but the `content` you send does not persist — it returns 200 with a body containing only `procedure_id` (`name` and `type` come back null), and the procedure lands with an empty body. You must follow it immediately with the draft PATCH below, resending the same `content`, or you will leave a named, empty procedure on the branch. Verify with the single-procedure GET after publishing, not with the POST response.
- EDIT or RENAME an existing procedure, and complete a create: `PATCH /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID/draft`. This writes to the branch draft. There is no direct `PATCH .../procedures/$PROCEDURE_ID` (no `/draft` suffix); that returns 405.
This endpoint requires `name`, `content`, AND `type` — all three. Omitting `type` fails with a 422 `{"type":"missing","loc":["body","type"]}` even on a plain body edit. This is the opposite of the create endpoint's behavior, so do not carry the create payload over unchanged.
- PUBLISH pending drafts: `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` with any partial-merge body publishes all pending procedure drafts on the branch into a new version. The simplest no-op body re-sends the branch's current prompt. Fetch it first scoped to this branch so you do not clobber branch-local differences, then PATCH it back unchanged.
Tell the user an edited or created procedure sits on the branch draft until this publish step runs, and that it only goes live on the production agent when the branch is merged.
### The full create sequence
Creating a procedure with a body takes three calls. Skipping the second leaves an empty procedure; skipping the third leaves it unreadable.
1. `POST .../procedures` with `name` + `content` → returns `procedure_id`. Content is NOT stored yet.
2. `PATCH .../procedures/$PROCEDURE_ID/draft` with `name` + `content` + `type` → stores the body on the draft. The response echoes the real `name` and content length; check that length matches what you sent.
3. `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` re-sending the branch's current prompt → publishes the draft into a new version.
Then verify with `GET .../procedures/$PROCEDURE_ID` (now 200, with content) and confirm `has_draft: false` in the list. Until step 3, that GET 404s and the body is only in the draft.
Build request bodies with a JSON serializer rather than hand-written strings — procedure content is multi-line markdown, and a raw newline in a hand-built payload fails with a `json_invalid` "Invalid control character" 422 that looks like a schema problem but is a quoting bug.
## 3. Content-shape gotchas that cause the failures
1. `content` is ONE string, not structured fields. For `free_form` it is a single markdown string folding the trigger and the steps into it. There are no separate `trigger` / `instructions` / `steps` parameters. Splitting them into separate args, or sending a nested object, is a top validation failure.
2. An edit draft replaces the whole body. Whatever `content` you send becomes the entire procedure; anything you omit is gone. Always start from the current content you read in step 1, modify it, and send the complete result. This is the number one cause of accidental data loss.
3. On the draft edit, `name`, `content`, and `type` are all required every time — even for a rename. To rename without touching the body, resend the existing content and type alongside the new name. To edit the body without renaming, resend the existing name and type.
4. `type` is REQUIRED on the draft edit — you cannot omit it to inherit the current type. Read the procedure's existing `type` from the list first and resend that exact value; sending the wrong one converts the procedure's form. Only `free_form` and `deterministic` are valid; an empty string or any other value fails. (`type` is genuinely optional on the create POST, where it defaults to `free_form`.)
5. Deterministic procedures: `content` must be a valid JSON string describing the flow. Malformed JSON, or free-form markdown while `type` is `deterministic`, fails. If unsure of the schema, read an existing deterministic procedure's content as a template, or see the `architect-structured-procedures` skill. Editing a deterministic procedure recompiles automatically, so keep the JSON well-formed. Compile explicitly via `POST /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/compile`.
6. `name` is capped at 200 characters. Keep it short; put the detail in `content`.
## 4. Recover from a failed call by error
- `schema_mismatch` (often no field detail): the content shape is wrong for the type. Re-check step 3. Usually you split trigger and steps into separate fields, sent an object instead of a string, or sent markdown for a `deterministic` procedure with broken JSON. Rebuild content as a single correctly-typed string; do not blind-retry the identical payload.
- `validation`: a concrete required arg is missing, usually `content`, `name`, or `type` on the draft edit, or `procedure_id`. Supply the named field, resending the existing value for anything you were not trying to change.
- `not_found` / 404 on the single-procedure GET: decide which case you are in. (a) The procedure exists but has no published content yet — it was just created, or only ever draft-edited. The GET 404s until a publish commits a body, while the list still shows the procedure. Confirm with the list; if it is there, this is expected and the fix is to publish, not to retry the read. (b) A genuinely wrong or stale id, or you are on a different branch than where it lives. Re-list procedures on the intended branch to get the current id, then retry. Do not carry a `procedure_id` across a branch switch.
- 422 `{"type":"missing","loc":["body","type"]}` on the draft edit: you omitted `type`. It is required here even though the create endpoint defaults it. Resend with the procedure's existing `type` from the list.
- 405 on `PATCH .../procedures/$PROCEDURE_ID`: you dropped the `/draft` suffix. Re-issue against `.../procedures/$PROCEDURE_ID/draft`.
## 5. Guardrail prompts are not in the procedure body
A procedure's guardrails are a separate first-class field on the procedure, not text inside `content`. If a user asks to edit "the guardrail prompt", do not look for it in the content string. Read the procedure to see its configured guardrails, tell the user plainly that the draft-edit path covers name, content, trigger, and type only, and do not smuggle guardrail text into `content` (it will not be enforced as a guardrail and silently changes the body instead).
## 6. Confirm and cross-reference
After a successful write, tell the user exactly what changed (created, edited, or renamed), on which branch, and that it sits on the branch draft until published and only goes live on merge. For branch and traffic-split workflow when staging a change for review before it goes live, see the branch and versioning guidance. If the user actually wants to change behavior via the prompt, LLM, voice, or tools rather than a procedure, that is a config change, not a procedure edit.
architect-migrate-from-competitor39.5 KB
---
name: architect-migrate-from-competitor
description: Use when the user wants to move a voice agent onto ElevenLabs Conversational AI from another platform — Retell AI, Vapi, or Bland — or hands over an export from one of them to work from. Fires on "migrate my Retell agent", "convert this Vapi assistant", "move from Bland to ElevenLabs", "port my agent to ElevenLabs", "here is my agent JSON", or when the user pastes a JSON blob containing response_engine, squadId, transportConfigurations, or pathway. Not for unrelated database, codebase, or framework moves.
---
# Migrate a voice agent from another platform to ElevenLabs
## What this skill is for
The user has an agent built on Retell, Vapi, or Bland and wants it running on ElevenLabs. This skill
owns the parts of that job specific to *importing*: identifying the export, handling credentials
inside it, mapping foreign constructs onto ElevenLabs ones, naming what genuinely has no equivalent,
and sequencing the work.
It does **not** re-teach ElevenLabs authoring. This plugin has ~30 skills that own that, and they
are more detailed and better maintained than a summary here would be. Your job is to translate and
to route.
**What counts as finished: an agent that demonstrably runs.** Not an optimal one — a conversion that
starts, resolves its variables, calls a tool successfully and reaches an end. Aim to get there in one
pass and one round of fixes, which means the checking in step 5 is part of the work rather than
something offered afterwards. The finish line is a run, never a re-read of your own config, and "I
converted everything" is not a status report.
Assume an authenticated EL workspace with `$API_KEY` set, as every skill in this plugin does.
## What to delegate, and to whom
Route each step rather than reimplementing it. Most migrations touch six or seven of these; read one
when you reach its step, not up front.
| Step | Skill |
|---|---|
| Read the target agent's current state | `architect-explore-agent` |
| Choose the model | `architect-llm-selection` |
| Port prompt text and first message | `architect-edit-string-fields` |
| Build the workflow graph, nodes, edges | `architect-edit-workflows` |
| Author step-by-step logic as procedures | `architect-structured-procedures`, `architect-manage-procedures` |
| Recreate tools | `architect-create-webhook-tool`, `architect-create-client-tool`, `code-tools` |
| Fix a tool that misbehaves | `architect-edit-existing-tool`, `architect-troubleshoot-tool-errors` |
| Post-call fields and evaluation criteria | `architect-post-call-data` |
| Remaining config, turn-taking, latency | `architect-update-config-safely`, `architect-turn-taking-latency` |
| Auth, guardrails, retention | `architect-secure-for-production` |
| Stage and ramp the rollout | `architect-branches-versions-merge`, `architect-schedule-launch` |
| Mock tools before any test run | `architect-mock-all-tools` |
| The smoke run that proves it works (step 5) | `architect-create-simulation-test` |
| Parity tests, CI | `-llm-test`, `-tool-test`, `architect-get-agent-ci-green` |
| Score the agent — only once asked, see step 6 | `agent-review`, then `agent-simplification` |
| Check it once real calls exist (step 6) | `architect-review-live-calls` |
**Do not create the agent through `architect-generate-agent`.** Its own documentation warns that its
flow regenerates the prompt instead of installing the one you just converted. Create the shell
directly with `POST /v1/convai/agents/create`, then install the converted content.
**Reading a sibling skill is not the same as routing to it, and the difference costs you at write
time.** The temptation is to skim the two or three you think you need and hand-roll the rest, which is
faster right up to the first rejected payload — a real migration reimplemented tool creation, post-call
data, model selection and all testing inline, and hit its schema surprises while writing to the live
workspace rather than while reading. Each of these skills owns field-level detail this one deliberately
does not restate. Before you report, account to yourself for which skill covered each step: a step you
cannot attribute is a step you improvised. That accounting is for you, not for the report — a customer
has no use for a list of internal skill names.
## Step 1 — identify the platform
Dispatch on a field that is always present, not on one that merely often is:
| Signature | Platform | Read next |
|---|---|---|
| `response_engine.type` | Retell | `reference/retell.md` |
| `squadId`, `members[]`, `transportConfigurations`, `model.toolIds` | Vapi | `reference/vapi.md` |
| `pathway`, `nodes[].type` with a `Default` node, or a Persona payload | Bland | `reference/bland.md` |
Ask the user which platform it came from if the blob is ambiguous. Never guess from prose in the
prompt text.
## Step 1b — audit the export before converting it
Cheap, mechanical, and it routinely finds live bugs in the source agent. A migration is the first time
anyone reads the whole graph at once, so do this before deciding what to build:
- **Dangling edges** — a condition with no destination. These are silent dead ends on paths that look
wired in the editor, and they land on happy paths as often as error paths.
- **Orphaned nodes** — nothing routes in, or nothing routes out. A whole feature that has never run.
- **Variables read but never assigned** — spoken in a prompt, written nowhere. The caller hears a
literal placeholder or a blank. **Scan code-node bodies, not just template syntax.** A variable read
as `dv.account_ref` inside a JavaScript node is invisible to a `{{...}}` search, so an export with a
dozen code nodes will under-report badly on the exact check that matters most. Rank what you find by
how many tools bind it: a variable bound on every tool that nothing supplies fails every tool call at
conversation start.
- **One identifier, two different definitions.** Check whether any two tool entries share an id or a
name while disagreeing on url or method. This is not tidiness — one export reused a single tool id
across a main-flow tool and a component tool pointing at *different environments*, so the agent read
from one and wrote to the other. Report it: an agent whose write paths leave the environment its name
claims is a finding the owner will want, and it decides how you key the id map in step 4.
- **Response paths that cannot resolve** — an assignment reading a field the API does not return, or a
path whose shape is wrong for the response (array indexing and object prefixes are the usual
offenders). These carry across silently and break the same way on the new platform.
- **Branch conditions with a null target** and no else path.
### Do the audit in full. Do not report it in full.
This is the part that decides whether a migration feels like arriving somewhere or like being handed a
bill. Run every check above at full depth — the accuracy is the point, and nothing else in this skill
finds these. Then write the findings to a file, and bring **two** things back to the user:
1. **What their agent does** — the intents it handles, the tools it calls, the shape it is in. A dozen
lines, and a diagram of the source graph where one helps. Name the phases using the source's own
node names so they can check you. This is the only evidence they have that you understood their
agent before you started rewriting it, and it is the cheapest trust available in the whole job.
2. **Only what you cannot proceed correctly without.** The test is mechanical: *would a wrong guess
here break the migration?* A variable that every tool binds and nothing supplies, a voice that does
not port, an environment that two halves of the graph disagree about. On a real 123-node export that
list was **two items**. It is normally one to three.
Everything else goes in the file with a one-line pointer. Not because it does not matter — it is the
most valuable thing in this step — but because of *when* it arrives. A defect ledger delivered before
anything runs reads as "here is how much work you have bought," and the findings are almost all
pre-existing bugs in an agent that has been in production for months. Lead with a wall of red and a
first-time user concludes the platform is the problem, closes the tab, and stays where they are.
So do not open with a count of defects. When nothing blocks the conversion, **say that first** — it is
usually true and it is the single most useful sentence you can write.
**Separate the three registers, in the file and in anything you say aloud.** They read completely
differently and mixing them is what makes a short list feel long:
- **Blocks the conversion** — needs an answer now. This is the only category that goes inline.
- **Already broken in the source** — live today, reproduced faithfully. Say so explicitly. "Your
current agent reads a placeholder to the caller here" is a gift; the same fact in an undifferentiated
list is an accusation.
- **Does not port** — the genuine gaps, from `reference/<platform>.md`.
Never silently fix a defect and never silently carry one across. Logging it in the file is not silence;
it is where the mapping log and gap list already live, from step 4.
### Check the source's claims against its own endpoints
Everything above reads the export. Some of the worst defects are not in it — a prompt asserts a fact,
the backend it calls disagrees, and no amount of reading the JSON can tell you. Say a prompt instructs
the agent to offer same-day slots up to 5 p.m. while the scheduling endpoint refuses anything after
4:30. Both halves look right on their own, and the agent offers a slot its own tool then rejects.
So where the export hands you an endpoint, verify the facts the prompt states about it. This means
calling a third party's production API, so it is fenced:
- **Ask the user first, naming the endpoints.** They may not own that backend, and it may be someone
else's live system.
- **Read-only endpoints only** — a lookup, an availability check, a status query. Never one that books,
pays, cancels, sends, or writes, whatever its name suggests.
- **Synthetic values only.** A probe is not the place for a real name, number, or booking reference.
- **Skip it if the endpoint needs a credential you recovered from the export.** Using that credential
to explore is a wider use than its owner authorised.
What you learn joins the findings. A prompt fact the backend contradicts is a caller-visible bug in the
source agent, and it is the kind nobody has reported, because it only shows up on the calls that hit it.
## Step 2 — credentials, before anything else touches the export
Exports embed live credentials: bearer tokens, Basic auth headers, API keys in tool headers, and
sometimes a secret inside a webhook URL's query string. Handle these first, because every later step
copies fields around.
1. Scan every tool for header values, URLs with embedded userinfo or query secrets, and any field
whose name contains `auth`, `token`, `key`, `secret`, or `bearer`.
2. For each real credential found, create an ElevenLabs secret and reference it — never paste the
literal value into a tool definition:
```
POST /v1/convai/secrets
{"type": "new", "name": "billing_api_token", "value": "<the value from the export>"}
```
The response carries a `secret_id`. Reference it in the tool's `request_headers` as
`{"secret_id": "<that id>"}`. Rotation is `PATCH /v1/convai/secrets/{secret_id}` with a body of
`{"type": "update", "name": ..., "value": ...}` — note the id is in the path, not the collection
endpoint you created against.
**Secrets are workspace-scoped, so check what else uses one before you touch it.**
`GET /v1/convai/secrets/{secret_id}/dependencies/{resource_type}` lists the resources depending on
a secret. Rotating or deleting a secret that other agents reference breaks them, and a migration is
exactly the context where someone tidies up a credential that looked unused.
3. Tell the user which credentials you found and that they are now live in a third system, so they
can decide whether to rotate at the origin. A credential sitting in an export file has usually
been shared more widely than its owner realises.
4. A value that is already a `{{template}}` reference is not a credential — do not resolve it to a
literal, which is how an agent ends up authenticating as the wrong principal for the rest of a call.
But do not carry the template across as a **string** either. **ElevenLabs substitutes nothing in a
header string** — a header given as text is sent exactly as written, so `"Basic {{vault_token}}"`
transmits those literal characters and the request is simply unauthenticated. It is a credential
reference that does nothing, and it fails on every call rather than visibly at build time. Use the
typed forms:
- a dynamic variable — `{"variable_name": "vault_token"}`
- a workspace secret — `{"secret_id": "<id>"}`
- an environment variable — `{"env_var_label": "<label>"}`
A locator supplies the **whole** header value, so it cannot be composed with a literal prefix. That
sounds like a problem for `Basic <token>` and is not: **put the scheme word inside the secret.**
Create the secret with value `Basic <token>` and reference it with `{"secret_id": ...}`. The header
is then complete on its own, and the credential never appears in the agent config. If you find
yourself wanting `"Basic {{vault_token}}"`, that is the sign to make a secret holding the whole
header value.
5. **Watch for one variable that plays two roles.** A single name is often both a *static seed*
credential used by early calls and a *per-user token* that a later tool overwrites mid-call. They
need different treatment and it is easy to see only one of them:
**The test is mechanical: does any tool write that variable during the call?** Look for a response
assignment on any tool whose target is the variable name — not just on the tool you are holding,
since the tool that issues a token is rarely the tool that uses it. Then:
- **Nothing writes it → static.** Its value lives in the agent's variable defaults, so a
`{"variable_name": ...}` locator would leave a live credential sitting in agent config in
plaintext. Make a secret holding the complete header value and reference it with
`{"secret_id": ...}`.
- **Some tool writes it → swapped.** Use `{"variable_name": ...}`. A secret cannot change mid-call,
and the locator leaks nothing because the value only ever exists at runtime.
Because the locator supplies the whole header, a swapped variable has to hold the whole header value
too. If the endpoint issuing that token returns a bare token, the assignment writing it cannot add
the scheme word — so either have that endpoint return the complete value, or confirm the receiving
API accepts a bare token. That is a short, concrete question for the customer, and it is worth asking
before you build rather than after the first unauthenticated call.
Handling only the swap leaves a placeholder in the agent's variable defaults, which invites someone
to paste a live credential into agent config in plaintext. Say explicitly which of the two you have
done.
Do not write the raw export, or any redacted copy containing real values, into a repo or a log.
## Step 3 — pick the shape
One decision to state and one to ask. What gets *shipped* has a sensible default and should not be put
to the user at all; what *shape* it takes is worth their judgement, and this is the moment for it —
after the audit, before anything is built.
### The deliverable is a default, not a question
**The rebuild ships. A faithful port gets built quietly alongside it as an oracle.** State that in one
line and move on. Do not open someone's migration by asking them to choose a migration philosophy —
they came to get their agent working, and the trade only means something to a person who already knows
what fidelity costs.
- **The rebuild is the agent.** The source is the specification, not the blueprint: pick the shape that
fits, fix what the audit found, hand them something callable.
- **The faithful port is built and then not mentioned again.** It mirrors the source construct for
construct, it is never deployed, and it is one converter pass now that the audit exists. It is the
oracle — what parity tests are written against, and what `agent-simplification` wants before it will
touch anything. One line in the report's file list is the whole of its visibility.
Two things change this, and both come from the user rather than from your reading of the export. If
they ask how they can tell the new agent matches, the port stops being silent — that is its purpose and
it is already built. If they say plainly that they are done with the source platform, skip it.
Findings are simply fixed in the rebuild and reproduced in the port. There is no per-finding question
to ask; the report says which ones you fixed.
### The shape is the question
This is the one decision worth putting to the user before you build, because it is the one they can
actually judge — they know their own call flow.
Four options, in increasing order of cost. **Recommend one** — with the reason, in a sentence they can
disagree with — rather than presenting a menu and waiting:
- **A single agent with a prompt.** Right for a source that is one big prompt with a few tools, and
right more often than the source's node count suggests.
- **An agent with procedures.** Right for step-by-step logic — verification sequences, intake,
scripted disclosures — where the order matters but the graph does not branch much.
- **An agent with a workflow.** Right when the source genuinely has distinct states with different
tools, models, or voices per state.
- **A mix.** Common and often correct: a prompt or a small workflow for the spine, procedures for the
sequences that hang off it. Do not treat the three above as exclusive.
Do not mirror the source's topology out of loyalty. Competitor editors encourage many small nodes, and
every node you carry over multiplies the configuration surface that can go wrong. Say what the shape
buys in terms of *their* agent — "your three queue components run the same algorithm, so they become one
paragraph and three procedures" — not in terms of node counts, which mean nothing to them. Every real
reduction comes from that kind of semantic merge; a mechanical conversion is roughly
node-count-neutral, which is exactly why the port is not the thing that ships.
## Step 3b — get the shapes right before anything is watching
Discovering the API's shape rules by being rejected by it is slow, and it burns the patience of whoever
is watching — a customer on a call, or someone who just uploaded their own export and is waiting. It is
avoidable: nearly every rejection below is a known shape rule, and the rest can be learned somewhere
harmless.
Three habits, in order of how much they save:
1. **Build every payload locally first, then write.** Do not alternate authoring and API calls
node-by-node. Compose the whole set — tools, nodes, edges — check it against the rules in this
skill and `reference/expressions.md`, and only then start writing. A shape mistake made once in a
local draft is one fix; the same mistake discovered on call forty is forty.
2. **Do any genuine schema discovery somewhere disposable.** If you truly do not know whether a field
is accepted, find out on a scratch agent rather than on the one being migrated. Ask which workspace
to use for it — a workspace holding production agents is not disposable, and tools and knowledge-base
documents are workspace-global, so a probe that creates one is not confined to your scratch agent.
Delete only what you created in this session, by id you captured at create time. Never learn on the
customer's agent, and never learn on a branch they are reading.
3. **Re-read a rejection before changing anything.** The messages below name a *symptom*, not the
field at fault, so the instinct to bisect the payload is wrong and expensive. One of these cost a
real migration fourteen identical failures.
| Rejection | What it actually means |
|---|---|
| `Input should be a valid dictionary or object to extract fields from` | An `expression` was passed as a **string**. Conditions are structured objects — see `reference/expressions.md`. |
| `Input should be a valid dictionary` | A `path_params_schema` was passed as an **array**. It is an object keyed by parameter name. |
| `Can only set one of: description, dynamic_variable, is_system_provided, constant_value, or is_omitted` | A property declares **two or more** value sources. Exactly one. |
| `Must set one of: description, dynamic_variable, ...` | A property declares **none**. Most often an array's `items` schema, which needs its own source. |
| `Invalid URL format` | A `{{...}}` source interpolation survived in a tool URL. Static host, single-brace path placeholder, `path_params_schema`. |
| `Duplicate edge found between X and Y` | An edge already exists for that node **pair**. Edges are keyed on the unordered pair, so the reverse transition becomes that edge's `backward_condition`. |
| `Input should be 'generate_immediately', 'wait_for_user' or 'auto'` | `entry_behavior` got something outside its enum. |
| `Input should be 'auto', 'force' or 'off'` | `pre_tool_speech` was given filler text. It is an enum, and it cannot carry a line — use `forced_tool_name` on an `override_agent` node instead. |
And one that never rejects at all, so no message will tell you: **`expression: {}`** is accepted and
silently becomes `boolean_literal` false, so the edge never fires.
And two on the **test** payloads, which is where people are still guessing after everything else has
gone in clean: `success_conditions` is a list of strings, and `simulation_scenario` is a plain string —
not the nested `simulated_user_config` object its name suggests. Both reject, and both only bite at
step 5, so they read as a broken smoke run rather than as two schema rules.
If something does go wrong in front of the customer, say what you are checking rather than narrating
surprise. "That field takes an object, fixing" is fine. A string of "oh, that didn't work" is what
makes a routine schema rule look like a broken migration.
## Step 4 — build in dependency order, and keep the id map
**Read `reference/traps.md` before you author anything here, and `reference/expressions.md` before the
first condition.** They are short and they are the failure list for exactly this step: every item in
them is a thing that is accepted by the API and wrong at runtime, so nothing later in this skill will
catch it for you. Read them now rather than after a rejection, because these do not reject.
**For the payload shapes themselves, read `reference/` — not a sample config you found in the
working repo.** A complete agent and a complete webhook tool live there, every field populated. The
sample you would otherwise grep for may be stale, may be someone's scratch file, and outside
ElevenLabs does not exist at all, which is how a skill passes internally and fails in the field.
Most of the order here is forced by what references what. Getting it wrong does not produce a clear
error; it produces a graph with references to things that do not exist yet.
1. **Create the agent shell.** Everything else needs its id. Do not use
`architect-generate-agent` (see above).
2. **Create every tool, and record the id each one comes back with.** This is the step people skip
the bookkeeping on, and it is unrecoverable later: a workflow's `tool` node references a tool by
**id**, not by name, and so does `additional_tool_ids` on an `override_agent` node. You cannot build
any node that calls a tool until its id exists.
3. **Create procedures through the branch-scoped flow** that `architect-manage-procedures` owns.
Passing them in the agent-create payload returns 200 and silently drops them.
4. **Build nodes, then edges.** Edges reference node ids, so the nodes have to exist. Order only the
conditional edges — the unconditional one is moved last for you.
5. **Then the rest of the config** — model, privacy, turn-taking — and only then step 5.
**Key the id map by something that is genuinely unique, which is neither the name nor the id alone.**
Exports break both keys, and each failure is silent:
- **Same name, different tools.** Routine. Once you have renamed one to satisfy uniqueness, a
name-keyed map cannot tell them apart and quietly points several nodes at one tool.
- **Same id, different definitions.** Worse, and the case step 1b tells you to look for. An id-keyed map
collapses the two, and half the graph then calls the wrong host — which nothing rejects, because both
hosts answer.
So key on the **definition**: the source id together with the fields that decide where a call lands,
url and method at minimum. Two entries are the same tool only when those agree. When they do not,
create two ElevenLabs tools and record in the mapping log why one source tool became two — that split
is exactly the kind of thing a reader will otherwise assume was a mistake.
### Leave a coverage proof, for the rebuild and the port alike
Two local files. They are the difference between output someone can audit and output they have to take
on trust, and they cost almost nothing because you are making each of these decisions anyway:
- **A mapping log** — one entry per source behaviour: what it was, where it landed, and why. In a
faithful port this is the fix log. In a rebuild it is the *only* thing that answers "did you drop
something?", which is the rebuild's one genuine weakness against a port.
- **A gap list** — every construct you could not map and what you did instead. This is also what the
user needs in order to sign off on the approximations, rather than discovering them later.
Write both to a scratch directory, never into a repo, and never let either hold a credential value or a
verbatim copy of the export.
### Then check your own output, not your intentions
Run a mechanical pass over what you actually emitted. Re-derive the things that are cheap to re-derive:
every edge's endpoints exist, every tool id referenced was really created, every node is reachable,
every variable a tool binds is assigned somewhere.
This is not ceremony. On one real migration that pass caught a bug in the converter its own author had
just written: where a node pair carried an edge in only one direction, the edge was emitted reversed and
unconditional, silently orphaning five nodes and one whole lookup path. Re-reading the code would not
have found it, and every validator in the chain accepted the payload. Report the check's result
alongside the payload; a clean result you can point at is worth more than an assurance.
## Model choice
Do not carry the source model across by name, and do not use any static table of compliant models.
`GET /v1/convai/llm/list` returns the models available to *that* workspace, filtered by the
deployment's data residency and the workspace's compliance entitlements. Call it, then hand the
choice to `architect-llm-selection`, which owns the region, HIPAA/PCI/ZRM and latency tradeoffs.
## Step 5 — prove it runs, before you tell anyone it is done
This is a step, not a formality at the end, and it is the difference between a migration that *looks*
finished and one that works. Every failure mode in `reference/traps.md` is silent: nothing rejects your
payload, the agent exists, the graph renders, and it is broken on the first real call. A migration
reported as complete without this has not been checked — it has been *hoped about*.
**The bar for "working" is five things, and all of them are observable:**
1. Every tool was created — no payload was rejected and quietly skipped.
2. The agent starts and delivers its first message.
3. Every dynamic variable a tool binds is actually supplied at conversation start.
4. At least one tool call executes and comes back successful.
5. At least one path reaches a terminal node.
Nothing there is about quality. A first migration should be judged on whether it runs at all; tuning
comes after, and confusing the two is how a broken agent gets shipped with a confident summary.
### Mock the tools for the first run, or you will transact against production
**Tool mocking defaults to off.** A simulation with default settings calls the customer's real
endpoints — so on the agent you just migrated, a smoke test books the appointment, charges the card,
and mutates the cart, for real. Nothing warns you.
So run it in two passes:
- **Pass 1 — everything mocked.** This is the cheap, side-effect-free run that proves the *graph* is
alive: first message fires, variables resolve, edges route, a terminal node is reached. It catches
the conversation-start binding trap outright — a variable a tool needs but nothing supplies fails the
run with a `missing_dynamic_variables` error naming it, which is exactly the check that is
impossible to pass by reading.
- **Pass 2 — read-only tools live, everything that mutates still mocked.** One real lookup is what
proves auth actually works, and auth is the thing most likely to be silently wrong after a
migration. Do not unmask a tool that books, pays, cancels, sends, or writes.
**Mocking everything is not the same as defining mocks, and the difference will waste your first run.**
When no mock matches a call the default is to raise, not to fall through to the real tool. That default
is the right one — an unmatched call is visible instead of quietly hitting production — but it means
"mock all" with no mock bodies makes every tool call fail, and you learn nothing about the graph. Give
each tool one mock with no parameter conditions, which makes it always match, and a `mock_result`
holding the minimal successful response shape that tool's consumers read. Where a response assignment
pulls a field out, the mock has to contain that field or the variable silently stays unset and you are
debugging the mock instead of the migration. A mock can also be marked as an error deliberately, which
is how you exercise a failure path once the happy one works.
`architect-mock-all-tools` owns the mocking setup and `architect-create-simulation-test` owns building
the run; route to them rather than hand-rolling. The current endpoints are
`POST /v1/convai/agent-testing/create` and `POST /v1/convai/agents/{agent_id}/run-tests` —
`simulate-conversation` still exists but is deprecated.
### Report what the run showed, not that you ran one
Give the five items above with their actual outcome, and name anything you could not verify. "Smoke
test passed" is not a result; "first message fired, three edges traversed, `fetch_profile` returned 200,
checkout still mocked and untested" is.
**Two rounds is the budget for fixing.** If a third would be needed, the failure is probably not a
conversion defect, and grinding on it is how a one-hour job becomes four. Stop fixing and work out
which of these it is: a success condition that does not match what the agent was asked to do, a mock
missing a field its consumer reads, or a real platform behaviour. Then say which, with the measurement,
and stop. **A behavioural gate that holds two runs in three is a finding, not a bug** — non-determinism
does not converge under repetition, so report the rate and what would make it absolute rather than
re-running until it looks better. Between rounds, re-run only what failed; run the whole suite once at
the end.
### The closing report
This is the artifact the migration is judged on, and its job is to leave someone confident enough to
move. Order matters more than content here — the same facts in the wrong order read as a warning.
1. **What it does, and what it became.** The phases and tools, and a before/after diagram of the source
graph against the new shape. This goes first because it is the only section that says *I understood
your agent*.
2. **That it runs.** The five checks with their outcomes.
3. **Mapped directly · Handled natively by ElevenLabs · Needs you.** Three groups, in that order. The
middle one matters more than it looks: a construct the source built by hand that this platform does
for you is a *win*, and filed as "difference" it reads as a loss. The third group is the only one
that asks anything of them, and it should be short.
4. **Decisions and why.** Every judgement call you made that they might have made differently — one
line each, with the reason and the alternative. A voice chosen as a placeholder, a model held at the
source's for comparability, an environment resolved one way. This is what makes the migration
reviewable rather than a black box, and it is the section people actually read.
5. **Not migrated, on purpose.** Everything net dropped: orphaned nodes nothing routed to, edges that
led nowhere, tools referenced by nothing, constructs with no equivalent. Say why for each, in a
clause. An unexplained absence is the thing that erodes trust later, when they go looking for a
feature and cannot find it — and most of these were already dead in the source, which is worth
stating plainly.
6. **The files, in one line each** — the audit findings, the mapping log, the gap list, the port.
7. **What to do next, and it is not a review.** If they have transcripts or recordings from the source
platform, this is the moment they are worth the most: real scenarios replayed against the new agent
find in an afternoon what invented ones will not find at all. Invite that. It is a better next step
than any score, and it is the one that turns a migration into their agent.
Do not lead with a defect count, do not grade the agent, and do not attach a scorecard. The step 1b
findings live in a file with a pointer, and `agent-review` is step 6's business and gated there.
## Before you call it done
**These five can only be answered by running something.** Do not substitute a re-read of your own
config for any of them.
- **A mocked smoke run passed** — first message delivered, variables resolved, a terminal node reached.
- **One read-only tool executed live and succeeded**, so auth is proven rather than assumed.
- **No `missing_dynamic_variables` error**, the only real proof that every bound variable is supplied
at conversation start.
- **Every tool was created**, counted against the source's tool list. A payload rejected and skipped
leaves a graph routing to nothing.
- **The self-check over your emitted output ran and you are reporting its result** (step 4) — not the
validator's, which passes payloads that are structurally wrong.
**Then the report, which is what the migration is actually judged on.**
- It leads with what the agent does and that it runs — not a defect count, not a score. The step 1b
findings are in a file with a pointer, and `agent-review` was not handed over unasked.
- It says what was dropped and why, and what you decided and why.
- Only what genuinely blocked the conversion was put to the user as a question. One to three is normal;
eleven means you moved the audit into the conversation.
- The mapping log and gap list exist and the user has them.
- Every credential you found is a secret reference and you said so — including any variable carrying
both a static seed and a swapped token. You **ran** the scan rather than reasoning over the tool list.
**Then the silent failures — and check them against the files, not against a summary.** Re-open
`reference/traps.md` and `reference/expressions.md` and read what you emitted against them: headers,
tool URLs, conditions, transfer nodes, node-scoped attachments. This file used to restate that list,
which meant two copies to keep in sync, and they drifted. One instruction is safer than a duplicate.
Parity tests are the one item that is genuinely follow-up rather than part of a working migration. Say
so plainly instead of implying the agent is test-covered when it has had one smoke run, and point at the
customer's own transcripts, which is where real coverage comes from.
## Step 6 — hand off, because "runs" is not "good"
This skill's finish line is deliberately low: a conversion that runs. Say that plainly and name who
answers the next question, rather than leaving the user thinking a smoke test was a quality review.
**But do not hand a scorecard to someone who has just arrived.** `agent-review` is an operator's tool.
Pointed at a fresh migration it returns a page of structural findings that read as a verdict on the
decision to switch — and it lands right after the audit and the test results, which is three documents
in a row telling a new user their agent is broken. That sequence is how a migration gets abandoned at
the last step. So gate it: run it when the user asks, or once they have confirmed the agent does what
they need. Not by default, and not in the same breath as the report.
1. **`agent-review`** — scores prompt quality and workflow structure, which is the axis this skill
defers on purpose. Right for an operator tuning an agent they already trust.
2. **`agent-simplification`** — acts on the structural findings. A faithful one-node-per-source-state
port is exactly its input, and it proves behaviour is preserved with a ground-truth suite rather
than asserting it.
3. **`architect-review-live-calls`** — once real traffic exists, finds what is actually failing in
production. Nothing before this point can tell you that.
**When the review does run, it is scoring the rebuild, so a bad score is a real finding.** You already
made the structural calls `agent-simplification` would have made, which means a pile of structural
findings says the shape was wrong — not that the review is being unfair to a migration. Do not reach
for "it's a migration" to excuse a score you earned.
The exception is if someone points a review at the **port**. That is supposed to score badly: it
mirrors the source's topology because that is what fidelity means, and the findings are a backlog
rather than defects. Refactoring them away before anyone has confirmed the new agent matches the old
one destroys the only thing the port was for.
## Reference files, and when each one is due
Four files matter to any one migration, and three of them are not optional reading. The point of the
split is that each is short enough to actually read at the moment it applies, rather than skimmed as a
wall up front.
| File | Read it | What it holds |
|---|---|---|
| `reference/<platform>.md` | at step 1, once you know the platform | construct-by-construct mapping, the genuine gaps, export shape. Read only yours — you do not need the other two. |
| `reference/traps.md` | **before step 4**, and again while authoring each surface | every way a migrated agent breaks silently, grouped by whether you are authoring tools, nodes, or config |
| `reference/expressions.md` | **before you write any condition** | what can be expressed at all, why a carried-over condition inverts, and the guard builders to emit |
| `reference/` | **before you write the first payload** | a complete agent and a complete webhook tool, every field populated, plus the jq recipes for anything they do not show |
That last one is the plugin's shared example directory, not this skill's. Go there rather than
looking for a sample config in the working repo: the one you find may be stale, may be a scratch file
from another session, and in an install outside ElevenLabs does not exist at all. Read it as a field
catalogue — it shows where each field lives and what shape it takes. Do not start a migration by
copying it, or you ship a maximalist config whose settings nobody chose.
`traps.md` and `expressions.md` are ElevenLabs-side, so they apply whichever platform you came from.
If you are about to author a condition or a tool and have not read them, you are about to make one of
the mistakes in them — they exist because a real migration hit these with the mapping table open.
Referenced files: 9
architect-migrate-to-procedures2.86 KB
--- name: architect-migrate-to-procedures description: Use when the user wants to migrate a system-prompt-driven agent to a procedure-driven setup, pull step-by-step logic out of a heavy system prompt, or "convert my prompt into procedures". --- # Migrate a system-prompt agent to procedures Guide the user through moving a prompt-heavy agent to a procedure-driven setup. Everything is branch-scoped over the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`). You need `$AGENT_ID` and the `$BRANCH_ID` you are working on. Procedure creates and edits follow the draft then publish model; see the `architect-manage-procedures` skill for the exact endpoint mechanics. ## 1. Check current agent state Read the agent with `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. Verify it is effectively system-prompt-only: an empty workflow with no custom nodes or edges, and no existing custom procedures beyond default system ones (list via `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures`). If it already has custom workflow nodes or existing procedures, flag this to the user, but do not stop the migration. ## 2. Analyze the system prompt for candidates Read the system prompt (`conversation_config.agent.prompt.prompt` off the agent GET). Identify good procedure candidates: sequential step-by-step instructions, and conditional logic or branching paths (for example, "if the user asks for X, do Y"). ## 3. Survey for clarity If any part of the system prompt is unclear or needs more detail to become a procedure, ask specific clarifying questions. Resolve all ambiguities before proceeding. ## 4. Propose suggested procedures Present a Markdown table with columns: - Name: a concise, human-readable name. - Trigger: the user intent or phrase that should activate it. - Brief content: a short summary of the steps or logic to capture. Reflect any clarifications from step 3 in the proposal. Ask the user if they are happy with the plan. ## 5. Create procedures one at a time Once the user approves and all questions are resolved, create the approved procedures one by one: `POST /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures` per procedure (this writes a draft). Do not create them all in a single turn. Create one, confirm it succeeded, then proceed to the next. For content shape and the free-form vs deterministic choice, see the `architect-manage-procedures` and `architect-structured-procedures` skills. Publish the pending drafts into a version when the set is complete via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. ## 6. Review and test Invite the user to review the new procedures and try them out. Once trimmed logic has moved into procedures, you can slim the system prompt with a `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` on the prompt field, keeping the changes on the branch until the user is ready to merge.
architect-mock-all-tools8.38 KB
---
name: architect-mock-all-tools
description: Use when a simulation test needs every tool call mocked so runs are deterministic and never hit live systems. Fires on "mock all tools", "mock the tools in this test", "the test says no mock matched", "tool returned an error in my test run", "stop my test calling the real API", or when a run fails because a tool errored rather than because the agent misbehaved. Also use before running a suite for a DOM/Architect-style agent.
---
# Mock all tools in a simulation test
A simulation test that hits live tools is not a test, it is a flaky integration run. Mock every tool the agent can reach so a failure means the agent behaved wrong, not that a dependency moved.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, `$BRANCH_ID`.
## The one mistake that breaks almost every mock
**Mocks match on tool ID, never on tool name.** An override keyed by a name, or by an ID that is not on *this* agent, is silently ignored. The tool then runs unmocked and you see this in the transcript:
```
Error: no mock matched for tool 'open_support_form'
```
That message means "your key did not match", not "mocking is off". It is the single largest cause of red simulation suites.
Two ways to get the key wrong, both common:
1. **A tool name as the key.** `{"open_support_form": [...]}` never matches. Verified directly: a test with seven name-keyed overrides had all seven miss.
2. **A stale ID.** Tool IDs are not stable across agents or workspaces. An ID copied from another agent, or left behind after a tool was recreated, resolves to nothing. A suite can carry a full set of overrides where *not one* ID exists on the agent under test.
Both fail the same silent way: config saves fine, `mocked_tool_ids` looks populated, every call still errors.
## 1. Resolve real tool IDs first
Never hand-write an ID. Derive the name-to-ID map from the agent you are testing.
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" \
-H "xi-api-key: $API_KEY" > agent.json
```
The IDs are at `conversation_config.agent.prompt.tool_ids`. Resolve each to a name via the workspace tool list:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/tools?page_size=100" -H "xi-api-key: $API_KEY"
```
Three traps when resolving:
- **Page to the end.** `has_more`/`next_cursor` paginate. A workspace can hold thousands of tools; stopping at the first page or an arbitrary cap makes real IDs look unresolvable and sends you chasing a bug that does not exist.
- **Duplicate names are normal.** The same tool name often exists under several IDs from earlier copies. Never pick by name from the workspace list. Map name to ID **only across the agent's own `tool_ids`**, so you pick the one this agent actually calls.
- **Not every callable tool is in `tool_ids`.** System tools (`start_procedure`, `end_procedure`, `load_memory_entry`, `transfer_to_agent`) and RAG run unmocked no matter what. Do not try to mock them; expect them live in the transcript and let success conditions tolerate them.
## 2. Write the mock config
`tool_mock_config` is an object, not a list:
```json
{
"tool_mock_config": {
"mocking_strategy": "all",
"fallback_strategy": "raise_error",
"mocked_tool_ids": ["tool_exampleplaceholder00000001"]
},
"tool_mock_overrides": {
"tool_exampleplaceholder00000001": [
{"parameter_conditions": [], "mock_result": "{\"result\":\"ok\"}", "is_error": false}
]
}
}
```
- `mocking_strategy` is exactly one of `all`, `selected`, `none`.
- **With correct IDs, `all` and `selected` both apply overrides.** Verified with paired runs that differed only in strategy: both passed and both returned the override payload. If an override is not landing, the cause is the key, not the strategy.
- `fallback_strategy: "raise_error"` is what surfaces an unmocked call as a loud error instead of a silent live call. Keep it: it is how you discover the tool you forgot.
- `mock_result` is a **JSON-encoded string**, not a nested object.
- `parameter_conditions: []` matches any arguments. Add conditions only when one tool must return different data per call:
```json
{"parameter_conditions": [{"path": "workspace_id", "eval": {"type": "exact", "expected_value": "ws_123"}}]}
```
## 3. Make mock data domain-plausible, not merely well-formed
A schema-valid mock carrying an implausible value still fails the test, because the agent reasons about the *content*. This is subtle and costs real debugging time.
A ticket test mocked a lookup with id `TRIAGE-4821`. The mock matched and returned cleanly, but the agent refused to act: the platform's real ticket ids are `agtqa_`-prefixed, so the agent judged the id malformed and redirected to a support form. Every success condition failed, and none of it was the agent's fault. Swapping in `agtqa_3901...` made the same test pass with no other change.
So mirror production shape in mock values:
- Match real ID prefixes and formats (`agent_`, `tool_`, `agtqa_`, `agtbrch_`).
- Return the fields the agent actually reads. An empty `{}` where the agent expects `status` or `title` reads as a broken record.
- Keep the mock consistent with the scenario prose. If the user says "open ticket X", the mock must contain ticket X.
The failure signature to recognize: mocks match (no `no mock matched`), yet the agent declines, redirects, or asks for clarification. That is implausible mock data, not a misbehaving agent.
## 4. Cover the full DOM surface for Architect-style agents
A DOM agent inspects the page before acting, so a "single tool" test in fact calls many. Mock the whole surface or the run dies on the first unmocked read:
`get_agent_config`, `list_tools`, `list_agents`, `get_workflow`, `get_page_text`, `get_interactive_elements`, `get_session_activity`, `execute_parallel_agents_tool`, `navigate_to_page`, `navigate_to_url`, `open_support_form`
Add the tools specific to the behavior under test (e.g. `create_webhook_tool`, `get_agent_ticket`). Mocking a superset is free; a missing one is a failed run.
Also seed `chat_history` with one agent turn. A DOM agent with empty history reliably dies at `Timed out after 60s waiting for the agent to produce its next turn` — a hang, not a verdict:
```json
{"role": "agent", "message": "Hi! I'm Architect...", "time_in_call_secs": 0,
"tool_calls": [], "tool_results": [], "interrupted": false,
"reasoning": [], "used_static_kb_document_ids": []}
```
Set `simulation_max_turns` to at least 6 for a DOM agent; too low returns inconclusive before the flow completes.
## 5. Verify the mocks actually applied
Never trust a green status alone. Read the transcript and confirm each expected tool returned **your** payload:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/test-invocations/$SUITE_ID" -H "xi-api-key: $API_KEY"
```
Walk `test_runs[].agent_responses[].tool_results[]` and check `is_error` and `result_value`. A test can pass while every mock misses — one verified run passed on the agent's own reasoning with all seven of its mocks erroring. Passing for the wrong reason is worse than failing, because it hides the broken config until the behavior changes.
Grep the transcript for `no mock matched` on every run. Any hit means a key is wrong, even when the suite is green.
## 6. Edit an existing test in place with PUT
`PUT /v1/convai/agent-testing/{test_id}` updates a simulation test in place and returns the updated body. The id is preserved, so every attachment and folder placement survives the edit. This is the correct way to repair mocks on an existing test.
```bash
curl -s -X PUT "https://api.elevenlabs.io/v1/convai/agent-testing/$TEST_ID" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" -d @fixed.json
```
Send the whole object, not a partial patch: `type`, `name`, `simulation_scenario`, `success_conditions`, `simulation_max_turns`, `dynamic_variables`, `chat_history`, `tool_mock_config`, `tool_mock_overrides`. Omitted fields are not preserved. `PATCH` is not supported and returns 405.
A 404 from `PUT` means the id is wrong, not that editing is unsupported. Re-read the id from the agent's attached list before concluding anything else.
Prefer `PUT` over delete-and-recreate. Recreating mints a new id, which silently detaches the test from the agent's suite; you then have to re-attach it, and any folder placement is lost. Only delete when you genuinely want the test gone, and confirm with the owner first for a test you did not create.
architect-post-call-data6.18 KB
---
name: architect-post-call-data
description: Use when the user wants to extract structured fields from finished calls (data collection) or define success/failure evaluation criteria for post-call analysis. Covers field types, enums, the independent-extraction rule, writing criteria as questions, UNKNOWN for early termination, and the 2000-character limit.
---
# Set up post-call data collection and evaluation criteria
Two related but separate post-call configs. Data collection answers "what facts were in the call?"; evaluation criteria answer "was the call good?". Both are written with `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`), branch-scoped, live on merge.
## Ground the design in real calls first
Before proposing fields or criteria, read the agent and a representative sample of its conversations (in parallel where independent):
- `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`: read the system prompt and `first_message` to see what the agent is positioned to gather, and read the existing `platform_settings.evaluation.criteria` and `conversation_config.platform_settings.data_collection` so you extend rather than clobber.
- `GET /v1/convai/conversations?agent_id=$AGENT_ID`: recent call volume and a sample to inspect.
- `GET /v1/convai/conversations/{cid}` on 2-3 of those: confirm the target data actually appears in transcripts, and note the exact phrasing the agent uses and where calls terminate early. These payloads carry customer PII; do not copy transcripts into other systems or logs, and respect zero-retention-mode accounts.
## Part 1: Data collection (structured extraction)
An independent post-call LLM extraction runs over the transcript for each field. Config lives at `conversation_config.platform_settings.data_collection`, an object keyed by field name. Each field has `type`, `description` (the extraction prompt), an optional `enum`, and value-source flags.
Design each field against these rules:
- `description` IS the extraction prompt. Be specific: "The caller's full legal name as confirmed during the call", not "Customer name". Keep it under ~500 characters; long descriptions dilute the instruction.
- Pick the type by the data: `boolean` for yes/no facts, `string` with an `enum` for categories (for example `call_outcome` with `["completed","declined","transferred"]`), `integer`/`number` for amounts. Using `string` for a yes/no fact yields "yes"/"no" text instead of a real boolean.
- Use `enum` liberally. Even a large enum (50+ values) gives stronger guidance than prose and prevents free-text drift. Cover every realistic outcome; a missing value comes back null.
- Fields are extracted independently, one LLM call each. A field's description cannot reference another field. "Only if coverage_active is yes..." will not work, because this extraction never sees another field's result. Bake any needed condition into this field's own description.
- Scalar values only, no comma-separated multi-value answers. If you need several values, split into a primary `enum` field plus a secondary free-text `string` field.
- Design for null. If the caller hangs up early, most fields will be null. That is expected, not an error; tell the user null means "not discussed".
Write it: `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` with body `{"conversation_config":{"platform_settings":{"data_collection":{...}}}}`. Merge with the existing fields you read above; do not replace the whole `data_collection` object unless the user explicitly wants a reset.
## Part 2: Evaluation criteria
Each criterion is a separate post-call LLM call scored SUCCESS / FAILURE / UNKNOWN. Its `conversation_goal_prompt` is capped at 2000 characters. Aim for 3-5 criteria, not an exhaustive list. Write each one against these rules:
- Write as a question, not an instruction. "Evaluate whether the agent collected all required fields", not "The agent should collect all fields". Criteria grade behavior; they do not tell the agent what to do (that belongs in the system prompt).
- Be specific about SUCCESS vs FAILURE, using the exact phrasing you saw in the transcripts: "Mark SUCCESS only if the agent read the disclosure starting with [exact phrase]".
- Always add UNKNOWN guidance for early termination. Every criterion that depends on a later phase needs a line like "If the call ended before reaching this phase, mark UNKNOWN rather than FAILURE." Without it, early hangups drag the success rate down for a step that never ran.
- Group related checks into one criterion. Instead of four disclosure criteria, combine: "Evaluate whether ALL of the following were read: (1)..., (2)..., (3)...". Fewer, richer criteria are cheaper (each is its own LLM call) and easier to read.
- Cover compliance (were required disclosures read?), completeness (were all fields collected?), accuracy (were facts stated correctly?), and boundaries (did the agent stay in scope?).
- Respect the 2000-character limit per criterion. Grouping is how you fit a thorough check; if a grouped criterion is still too long, split along a natural seam rather than trimming the specificity that makes it useful.
- Do not make it brittle. Criteria that fail on valid conversations make the success rate misleading.
Write it: `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` with body `{"platform_settings":{"evaluation":{"criteria":[...]}}}`. Send the FULL array; the write replaces the criteria list, so include the criteria you read above plus the new ones.
## Verify
Criteria only mean something once conversations are scored against them. Have the user run a simulation test or wait for a few live calls, then re-open graded conversations with `GET /v1/convai/conversations/{cid}` and confirm the SUCCESS / FAILURE / UNKNOWN verdicts match your judgment. If a criterion fails on a call you would call good, it is too strict; loosen the wording or add the missing UNKNOWN branch.
## Related
- If the user wants a value known BEFORE the call (for example an account id passed in at session start), that is a dynamic variable, not a data collection field.
- If fields later come back null or wrong, re-check description specificity, type match, enum coverage, and whether the data is even in the transcript.
architect-review-live-calls12.9 KB
---
name: architect-review-live-calls
description: "Review an agent's recent finished conversations to find what is actually going wrong in production, then route each finding to a fix. Use when the user says 'we just launched this agent, how is it doing', 'review the last 3 days of calls', 'check how the agent is performing', 'monitor agent {id}', 'what are callers running into', 'why are calls failing', or asks for a post-launch health check or a weekly call review."
---
# Review an agent's live calls
Every other conversation read in this plugin is downstream of a problem someone already found — a QA ticket, a reported tool error, one bad call to turn into a test. This skill is the entry point that *surfaces* the problem: start from recent calls, end with clustered findings routed to a fix.
This skill owns the list→paginate→transcript fetch pattern. When another skill needs a sample of real calls, point at this one rather than reimplementing pagination and the PII rules.
This is a review of **finished** calls. Live in-flight audio over `wss://.../v1/convai/conversations/{id}/monitor` is a different thing and is not covered here.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY` and `$AGENT_ID`. For an EU/IN/SG residency workspace the host is `https://api.<eu|in|sg>.residency.elevenlabs.io` and the key is workspace-specific. Sending a residency key to the global host fails with a clear message — `Invalid API key: The API key used is for a data residency stack, however this server is a global server` — so if you see that, switch hosts rather than doubting the key.
## Two modes — pick one before fetching
Baselines live at a fixed path so a later run — even a different session — can find one without being told where to look: `.architect-live-calls/<agent_id>-baseline.json`, relative to the repo root. Check for that file before you fetch anything.
- **Launch sweep** (default; use when the agent went live within roughly the last two weeks, or when no baseline file exists at that path for this agent). Question: *is this thing working at all?* Did calls complete, did tools fire, where do callers drop, is the first message landing. Small N, no baseline, no QA tickets to route into — findings go straight to a fix.
- **Drift check** (use when the baseline file exists for this agent). Question: *what changed?* Compare this window's aggregates against the baseline and lead with deltas. Without a baseline a rate is not a finding — "12% of calls failed" means nothing on its own, so if the user asks for a drift check and no baseline file exists, run a launch sweep instead and write the baseline for next time.
State which mode you are running in one line before you start, and name the baseline path when you are running a drift check or writing a fresh baseline.
## What the list endpoint gives you — and what it withholds
`GET /v1/convai/conversations` returns `{conversations, next_cursor, has_more}`. Each entry carries enough to run the whole aggregate pass without touching a single transcript:
`conversation_id`, `agent_id`, `agent_name`, `branch_id`, `version_id`, `start_time_unix_secs`, `call_duration_secs`, `message_count`, `status`, **`termination_reason`**, `call_successful` (`success` | `failure` | `unknown`), `call_success_score`, **`tool_names`**, `main_language`, `direction`, `conversation_initiation_source`, `sentiment_analysis`, `tag_ids`, `call_summary_title`.
Query params that are honored: `agent_id`, `page_size` (**max 100** — larger 422s), `call_start_after_unix`, `call_start_before_unix`, `call_successful` (enum-validated), `user_id`, and `cursor` (pass back `next_cursor`, loop while `has_more`).
**`summary_mode=include` is the highest-value flag here.** It defaults to `exclude`, and turning it on populates `transcript_summary` on every entry — a one-line account of each call. Cluster from those summaries *before* you open any transcript; it is far cheaper and far less PII-exposing than reading calls to find out what they were about.
**What the list does NOT give you:** `evaluation_criteria_results` and `data_collection_results` come back `null` in list responses *even with `summary_mode=include`*. They only populate on the single-conversation `GET /v1/convai/conversations/{conversation_id}`, under `.analysis`. So per-criterion pass rates and extraction fill rates are **sample statistics from step 3, never window-wide rates** — say so wherever you report them.
## 1. Set the window and fetch
Ask for the window if the user did not give one — they will usually say it in passing ("last 3 days", "since Monday", "this week"). Default to 7 days for a launch sweep.
```bash
curl -s "https://api.elevenlabs.io/v1/convai/conversations?agent_id=$AGENT_ID\
&call_start_after_unix=$SINCE&summary_mode=include&page_size=100" \
-H "xi-api-key: $API_KEY"
```
Filter server-side with `call_start_after_unix` — do not pull the agent's whole history and trim client-side. Loop on `cursor`/`has_more` until you pass the window edge.
If the window is empty, say so and stop. An agent with no calls is not a passing health check.
Alongside the window fetch, pull the agent's configured tool set once — it is what step 2's dead-tool check compares `tool_names` against, and the list endpoint has no way to derive it on its own: `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` for `conversation_config.agent.prompt.tool_ids` (plus each node's `additional_tool_ids` for a workflow agent), then `GET /v1/convai/tools` to resolve those ids to names. Skip this fetch, and skip the dead-tool check, if the user only wants a scoped read on a single aggregate.
## 2. Aggregate pass — count before you read
Build a small table from the list payload alone:
- **Volume and completion** — total calls, and the `status` / `call_successful` split.
- **`termination_reason`, grouped and counted.** Usually the single most useful number in the report — a spike in one reason is the finding.
- **Duration and `message_count` distribution** — flag both tails. A cluster of sub-two-turn calls usually means the greeting, the connection, or the language is wrong; a cluster of very long calls usually means the agent cannot close or is looping.
- **`tool_names` coverage** — union the arrays and compare against the configured tool names from step 1. A tool that appears in zero calls is either dead weight or an instruction the LLM never acts on, and that is invisible from the config alone.
- **`main_language` split** — unexpected languages mean detection or ASR problems, not multilingual success.
- **`call_success_score` and `sentiment_analysis` distribution** — useful as a tiebreak when `call_successful` is mostly `unknown`.
Then read the `transcript_summary` lines end to end and group them. This is the cheapest clustering signal available and it comes free with the same request.
## 3. Deep-read only the outliers
Pick targets from the aggregates, do not sample randomly: the dominant `termination_reason` bucket, `call_successful=failure` (filter it server-side), the tails of the duration distribution, and any summary cluster whose mechanism you cannot infer.
Cap it. Default 10 transcripts, 20 if the user asks for a thorough review. On `GET /v1/convai/conversations/{conversation_id}` the fields that earn the fetch are:
- **`.analysis`** — `evaluation_criteria_results` and `data_collection_results` (the only place they exist), plus `transcript_summary` and `call_success_score`.
- **`.metadata`** — `termination_reason`, `error`, `cost`, `phone_call`, `authorization_method`, `rag_usage`, `features_usage`.
- **`.transcript[]`** — per turn: `tool_calls` and `tool_results` (the failing call and its error verbatim), `interrupted`, `triggered_guardrails`, `conversation_turn_metrics` (latency), `rag_retrieval_info`, `original_message` vs `message`, and `role`.
**If you hit the cap, say what you dropped and why** — "read 10 of 34 calls in the dominant termination bucket" is honest; silently reading 10 and reporting as if you covered the window reads as full coverage when it isn't.
## 4. Cluster by root cause
Group by mechanism, not by surface wording — same broken tool, same prompt gap, same missing procedure branch, same ASR failure. **Corroborate every cluster against at least two conversations before you trust its root-cause hypothesis.** One transcript supports a story; two support a cause. A cluster of one is a suspicion — label it as one.
Rank by frequency × severity. A rare hard failure can outrank a common cosmetic one; say which you are doing when it is not obvious.
## 5. Route each cluster to a fix
Findings are worthless as a list. Hand each cluster to the skill that fixes that class of problem:
| Cluster looks like | Route to |
| --- | --- |
| Tool errored, timed out, or never appears in `tool_names` | `architect-troubleshoot-tool-errors` |
| Agent said the wrong thing, missing rule, weak greeting | `architect-edit-string-fields` |
| Missing or mis-branching step-by-step logic | `architect-manage-procedures` / `architect-structured-procedures` |
| Wrong routing, dead-end node, bad edge condition | `architect-edit-workflows` |
| Criteria that never fire, or extraction fields always null | `architect-post-call-data` |
| `interrupted` turns, slow `conversation_turn_metrics`, callers cut off | `architect-turn-taking-latency` |
| Unexpected `main_language`, misheard input | `architect-update-config-safely` (ASR and language config) |
| Broad structural problems across many clusters | `agent-review` for a full config audit |
Two rules on the handoff: fixes land on a branch, never straight on main (`architect-branches-versions-merge`), and **every confirmed cluster becomes a regression test before the fix ships** — `architect-create-llm-test` for a single bad reply, `architect-create-tool-test` for a wrong or missing tool call, `architect-create-simulation-test` for a multi-turn flow that fell apart. A production bug that ships a fix without a test comes back.
## 6. Write the baseline
Write to `.architect-live-calls/<agent_id>-baseline.json`, relative to the repo root — the same path checked under "Two modes" above to decide launch sweep vs drift check. Create the directory if it does not exist. One file per agent; overwrite it on every run so the next run always diffs against the most recent window, not a stale one.
Persist: window bounds, total calls, `termination_reason` counts, the `call_successful` split, tools seen, and the cluster labels with their counts.
Mark each number as window-wide (anything from step 2) or sample-derived (anything from step 3, including all per-criterion and fill-rate figures). A baseline that silently mixes the two produces fake deltas on the next run.
**Aggregate counts and conversation ids only.** No transcript text, no `transcript_summary` strings, no caller names, phone numbers, emails, or other caller identifiers — the baseline outlives the session and is easy to commit by accident. Check the repo's `.gitignore` for `.architect-live-calls/` and add an entry if it is missing before you write the file.
## Output format
Short. The report is roughly 25–50 lines:
1. One line: mode, window, call count, and the single most important number.
2. Aggregate table from step 2.
3. **Clusters** — one line each: what happens, how many calls, corroborating conversation ids, root cause, where it routes. Mark single-call clusters as unconfirmed.
4. **Top 3 to fix first**, ordered.
5. For a drift check, a deltas-vs-baseline line before the table.
No per-conversation walkthroughs, no praise section, no restating what the agent is configured to do. If a cluster needs more than two lines, the extra belongs in the skill you route it to. Produce the long version only if asked.
## Rules
- **Read-only.** This skill only issues GETs — against the conversations endpoints, and against the agent/tools endpoints to resolve configured tool names for the dead-tool check. It never deletes a conversation and never edits the agent — every config change goes through the sibling skills in step 5, which branch first.
- **PII.** Conversation payloads, summaries, and analysis carry customer PII, and calls may have audio (`has_audio`). Do not copy transcripts or summaries into other systems, logs, artifacts, committed files, or scratch files. Quote at most a short redacted fragment when it is needed to make a finding concrete, with names and numbers replaced.
- **Secrets.** Read `$API_KEY` from the environment. Never hardcode it in a generated script or in any file this skill writes, and never echo it to stdout or into the report.
- **Zero-retention workspaces.** Transcripts are unavailable by design, and `hiding_reason` on a conversation marks one you cannot read. Do not treat either as an error — run the aggregate pass on metadata alone, skip step 3, and tell the user which findings were unreachable without transcripts.
- **Non-actionable is a valid verdict.** A window where everything worked is a real result. Say so in a line and stop, rather than manufacturing findings to fill the report.
architect-schedule-launch6.02 KB
---
name: architect-schedule-launch
description: Use when preparing an agent change now but applying it later at a specific moment (product launch, marketing go-live, embargo, scheduled announcement). Fires on "queue up these changes and publish at launch", "prepare this now but don't make it live until tomorrow", "stage a knowledge-base update for the launch", "schedule a change", or "have this ready to merge when we go live".
---
# Schedule a change for a launch
The user wants to make changes now that must not affect live traffic until a later go-live moment. Model this as a branch: all work happens on a non-live branch, main keeps serving the current config, and the change goes live only when the user merges (or ramps traffic to) the branch at launch time. There is no timer. Scheduling here means staging on a branch and holding the merge until the user is ready. Be explicit with the user that nothing publishes automatically; they press merge or ramp at go-live.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY` and `$AGENT_ID`. For the exact branch mechanics, error recovery, and traffic-split rules, follow the branches, versions, and merging skill for every branch write.
## 1. Confirm the shape of the change
Ask or infer two things before touching anything:
- **What changes** - a prompt/config edit, a new tool, a knowledge-base document add or rescrape, or a fact the agent should know. A KB add or rescrape counts and is a common launch case.
- **When it goes live** - the launch moment. Confirm the user will trigger the merge or ramp themselves; you are only staging.
If the change would be fine to ship immediately, say so and skip the branch. Scheduling only earns its complexity when the change must stay dark until a specific moment.
## 2. Read current state first
Before any write, read the branch list (`GET /v1/convai/agents/$AGENT_ID/branches`) for exact ids, `current_live_percentage`, `protection_status`, `parent_branch_id`, and `draft_exists`, and get the currently active branch config (`GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`). You need main's exact id to branch off it and to know whether main is `admin_perms_required`, which gates the eventual merge.
## 3. Create the launch branch off main
Create the branch with a clear launch-named `name` (e.g. `elevenmusic-audio-references-launch`) and `parent_version_id` set to main's HEAD version (main's `.most_recent_versions[0].id` from the branch list):
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"name": "<launch-name>", "description": "Staged for launch", "parent_version_id": "<main-head-version-id>"}'
```
The new branch starts at 0% live traffic, so creation alone sends zero live traffic. That is exactly what "queue it" means. Confirm to the user the branch exists and is at 0%.
## 4. Make the change on the launch branch
Scope every edit to the launch branch with `?branch_id=<new-branch-id>` so it lands on the launch branch and never on main.
- **Config / prompt / tool edits** -> apply via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=<new-branch-id>` with a partial body (see the update-config-safely skill). Never spread edits across branches.
- **Knowledge-base add or rescrape** -> create or update the KB document (`POST /v1/convai/knowledge-base/text` or `POST /v1/convai/knowledge-base/url`) and attach it while working on the launch branch. The KB reference is captured in the branch's config snapshot, so it stays off main until merge. Confirm the document shows the updated content on the branch.
- **A fact the agent should recall** -> there is no memory endpoint, and even if there were, agent memory is agent-global, not per-branch, so a fact added "on the launch branch" would go live on main and every branch immediately. For anything that must stay dark until go-live, put the fact on the launch branch as a system-prompt line (via a branch-scoped PATCH) or as a knowledge-base document. Those are part of the branched config and only ship on merge.
## 5. Optionally publish and validate on the branch (does not touch main)
Because the branch is at 0% traffic, publishing a version on it makes the change testable without exposing it to live users. Offer whichever staging style fits:
- **Hold-and-merge (simplest):** leave the change as staged work on the branch. At launch, merge.
- **Publish-then-merge (testable):** commit the change on the branch (a branch-scoped PATCH returns a new version_id), let the user or a test suite validate it, optionally send a small traffic split for a pre-launch A/B, then merge or ramp to 100% at go-live.
## 6. At launch: merge or ramp (the user triggers this)
Do not do this step until the user says the launch is live. Then, per the branches skill:
- **Merge** the launch branch into main with `POST /v1/convai/agents/$AGENT_ID/branches/<launch-branch-id>/merge`. Always recommend a merge preview first to show the diff. If main is `admin_perms_required` and the user is not an admin, the merge fails with a permission error. Flag this now, in step 2, not at launch, and route them to an admin or to the traffic-split path.
- **Or ramp traffic** with a traffic-split change if they want a gradual rollout (10 -> 50 -> 100) rather than a hard cutover. The split must total exactly 100 across active branches in a single call.
## 7. Confirm and leave a clear handoff
Summarize in one place: the branch name and id, what is staged on it, that main is unchanged and still live, and the exact action the user takes at launch ("merge the <name> branch" or "ramp traffic to 100%"). If they queued several launches, keep one branch per launch so each can merge independently.
## Multiple people working at once
Give each person their own branch off main so their work does not collide. When two branches change the same field and both merge, the platform shows a conflict warning and keeps the more recent value. Call this out so the user reviews the merge preview rather than trusting a silent auto-resolve.
architect-secure-for-production6.87 KB
--- name: architect-secure-for-production description: "Use when a customer asks how to make their agent safe, secure, or production-ready, or how to roll it out without risk. Walks through the layered safety model across authentication, guardrails, testing, gradual rollout, real-time monitoring, and privacy." --- # Secure an agent for production Walk the customer through ElevenLabs' layered safety model. The guiding principle: protections at every stage, before the agent goes live, while it is running, and after each conversation. This is the same lifecycle behind ElevenLabs' AIUC-1 certification (the first safety/security/reliability standard purpose-built for AI agents, which also underpins agent insurance eligibility). Do not quote certification percentages, coverage terms, or pricing from memory; point the customer to their account team or current docs for anything commercial. Most of the settings below are configured in the agent's Security, Advanced, and Privacy config and can be inspected or set over the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`) with `GET`/`PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. Read the current config before advising. ## Stage 1: control who can reach the agent (authentication) Restrict access first. Configure one method per agent: - Signed URLs, recommended for client-side apps. Your server requests a temporary signed WebSocket URL using your API key, the client connects with it, and it expires after fifteen minutes. This authenticates sessions without exposing the API key client-side. - Allowlists, to restrict connections to approved hostnames (up to ten, exact-match, so add subdomains separately). Best for hostname-based access control. Do not configure both on one agent; pick the one matching the deployment model. The signed-URL mechanism confirms the request came from an authorized source; to restrict to specific users, authenticate them in your own app before requesting the signed URL. Never expose the API key in client code. ## Stage 2: shape and constrain behavior (guardrails) Guardrails protect at three levels: - System prompt hardening plus the Focus guardrail, the foundation. Put the most critical rules under a `# Guardrails` heading in the system prompt (models attend to it specifically), and enable the Focus guardrail to keep the agent on-topic across long conversations. - Manipulation guardrail, which validates user input, detecting prompt injection and instruction-override attempts and terminating risky conversations before the agent responds. - Content and Custom guardrails, which independently validate the agent's replies before delivery. Content blocks inappropriate material; Custom lets the customer define business-specific rules in plain language (for example, "block specific financial advice", "no refunds unless eligibility is confirmed"). For the most critical rules, put them in both the system prompt and a custom guardrail: defense in depth, so a response validator catches drift even if the model slips. Tradeoffs to convey: streaming mode adds no latency but may emit a little output before a block (recommended for voice); blocking waits for validation (~200-500ms, recommended for text) and is the only mode that supports retry as an exit strategy. Focus/Manipulation/Content are included; Custom guardrails incur usage-based LLM cost per response. Guardrails are currently in alpha; advise validating and monitoring logs as the feature evolves. ## Stage 3: prove it works before going live (testing) Do not ship on intuition. The testing framework supports: - Scenario tests, single-turn, evaluating a response against plain-language success criteria with success/failure examples. - Tool-call tests, verifying the agent calls the right tool with correct parameters (exact match, regex, or LLM evaluation). Essential for high-stakes actions like transfers. - Simulation tests, full multi-turn conversations against a simulated user, with tool mocking so live systems are not hit. Recommend turning real failed conversations into tests, testing prompt-injection attempts explicitly, and using probabilistic testing (run a test 3x/5x/15x via `repeat_count` on the run) to get a pass rate rather than a single pass before shipping a change. Integrate into CI/CD. Over REST, create tests with `POST /v1/convai/agent-testing/create`, attach with `POST /v1/convai/agents/$AGENT_ID/testing/attach-test`, and run with `POST /v1/convai/agents/$AGENT_ID/run-tests` (which takes `repeat_count`). ## Stage 4: roll out gradually When rolling out an agent for the first time, and when shipping new behavior to a live agent, do it gradually behind a feature flag. Do not flip the new agent on for 100% of traffic at once. Gate the new agent or behavior and ramp exposure (small cohort, then wider, then full), watching monitoring and guardrail logs at each step, with the flag as instant rollback. Versioning and Experiments (A/B on production traffic) support comparing the new configuration against the current one with data before fully committing. Branches and traffic splits over the API let you stage a new version and shift traffic to it gradually rather than merging straight to full exposure. ## Stage 5: watch it live (real-time monitoring, enterprise) For high-stakes deployments, real-time monitoring (enterprise-only) streams live conversation events over a WebSocket and lets an operator send control commands mid-call: end the call, transfer to a number, inject a contextual update, or trigger human takeover in chat. Enable it in Advanced settings; it needs `ElevenLabs Agents Write` scope and EDITOR workspace access. Use it for QA, human escalation, and call-center oversight. Limits: text/metadata only (no audio), roughly the last 100 events cached, connect only after the conversation starts. ## Stage 6: control what is retained (privacy) Match data handling to the customer's compliance needs: - Retention: how long transcripts and audio are stored (down to 0 days for immediate deletion). - Audio saving: whether call recordings are kept at all. - Conversation history redaction (enterprise): strips sensitive entities from stored transcripts (placeholders) and audio (bleeps); Zero Retention Mode (enterprise) for the strictest cases. Give conceptual guidance only on regulatory specifics. For HIPAA/GDPR retention periods, point the customer to current docs and their own compliance/legal team rather than asserting requirements. ## How to walk a customer through this Diagnose where they are. "Just built it, how do I make it safe?" points to all six stages in order. "Worried about it saying the wrong thing" points to guardrails (Stage 2) plus testing (Stage 3). "How do I launch safely?" points to gradual rollout (Stage 4) plus monitoring (Stage 5). Ground every step in its config location, and route commercial questions (insurance, AIUC-1 scope, custom guardrail pricing) to the account team rather than answering from memory.
architect-structured-procedures5.62 KB
---
name: architect-structured-procedures
description: Use when you want to create, rewrite, inspect, or fix a structured / deterministic / flow-style procedure, convert a procedure between free-form and structured, or when a deterministic procedure fails validation or the compile step rejects the JSON. The authoritative JSON schema is inline here.
---
# Write or convert a structured (deterministic) procedure
Both procedure forms are first-class. Pick deliberately:
- `free_form`: markdown with optional frontmatter and prose steps. The model reads it and applies judgement. Right when steps are guidance or need discretion. This is the default and usually the correct answer.
- `deterministic`: JSON compiled into a real state machine. Steps run in order, exactly as written. Right when a tool sequence must be guaranteed (compliance, ticket hygiene, ordering that must not vary between runs).
Never convert one form to the other unless the user asked. Converting a working free-form procedure to deterministic trades away flexibility; say which capability is being traded, in one sentence, when you convert.
Everything is branch-scoped over the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`). You need `$AGENT_ID` and `$BRANCH_ID`. Create and edit follow the draft then publish model (see the `architect-manage-procedures` skill): a create posts a draft, an edit patches the draft, and a `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` publishes pending drafts into a new version.
## Do not guess the schema
There is no `tool_code` step type and no nested `tool_code` object. Use the schema below, or read an existing deterministic procedure's `content` as a template (list procedures via `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures`, which returns `type` and full content for every procedure). Validate the JSON before publishing with `POST /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/compile`.
## Authoritative content schema
`content` is a JSON string with exactly two top-level keys. `steps` must be non-empty.
```json
{ "trigger": "<when this procedure fires, or empty string>", "steps": [ ... ] }
```
Every step carries a `type` discriminator. The complete set:
```
{ "type": "ask", "instruction": "<what to ask the user>" }
{ "type": "tell", "instruction": "<what to convey, model phrases it>" }
{ "type": "say", "message": "<verbatim message>" }
{ "type": "tool_call", "tool_id": "tool_...", "tool_name": "<name>",
"instruction": "<optional>", "on_failure": <handler|null> }
{ "type": "system_tool", "system_tool_name": "end_call" }
{ "type": "branch", "branches": [ <arm>, ... ], "fallback": [ <substep>, ... ] }
```
- An arm is `{ "condition": <condition>, "steps": [ <substep>, ... ] }` with non-empty `steps`.
- A condition is `{ "type": "llm", "condition": "<natural language>" }` or `{ "type": "expression", "expression": <ASTNode> }`.
- A substep may be `ask` / `tell` / `say` / `tool_call` / `system_tool`. A branch cannot nest another branch.
- `on_failure` is `{ "branches": [ { "condition": ..., "steps": [...] } ], "fallback": [ ... ] }`, where `fallback` is mandatory and non-empty, and its substeps may be `ask` / `tell` / `say` / `retry`. `retry` is `{ "type": "retry", "max_retries": 1-3 }`.
## Validation rules that reject your JSON
1. `ask` / `tell` need a non-empty `instruction`; `say` needs a non-empty `message`.
2. All conditions in one branch (or one failure handler) must be the same type; never mix `llm` and `expression`.
3. A `retry` step must be the last step in its failure branch.
4. `end_call` must be the last step wherever it appears.
5. No two consecutive `branch` steps; put a non-branch step between them.
6. An `expression` condition may not come directly after an `ask`. Use an `llm` condition to branch on a user reply, or place the expression after a `tool_call`.
7. `tool_id` must be a real tool on this agent. Get it from `GET /v1/convai/tools` (or the agent's `tool_ids`), never invent one, and make `tool_name` match.
## Converting free-form to deterministic
The failure mode to avoid: writing the JSON but leaving the procedure `free_form`, so the JSON lands as raw text in the markdown body and the user's prose procedure is replaced by a wall of JSON.
1. Read the current `content`, `name`, and `type` via `GET /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID`.
2. List tools via `GET /v1/convai/tools` to resolve every tool the prose mentions into a real `tool_id`, before writing any JSON.
3. Map prose to steps, preserving tool order exactly: "ask the user X" becomes `ask`, "tell them Y" becomes `tell`, a fixed sentence becomes `say`, "call tool T" becomes `tool_call`, "if Z then" becomes `branch`.
4. `PATCH /v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID/draft` with `type: "deterministic"` set explicitly, plus `name` and the JSON string as `content`.
Setting `type` is the whole conversion. An omitted `type` resolves to the procedure's existing type, so omitting it on a free-form procedure keeps it free-form. Editing the content string alone always writes back the existing type and therefore cannot convert a procedure; only a draft edit with an explicit `type` can.
5. Compile via `POST .../procedures/compile`. A compile error names the offending path (e.g. `steps[2].tool_id`); fix that field and resend once.
6. Publish the draft via `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. Confirm which form it is now, and that the edit sits on the branch draft until published and only ships on merge.
Never silently drop a step you could not map to a step type. Tell the user instead.
architect-troubleshoot-tool-errors3.92 KB
---
name: architect-troubleshoot-tool-errors
description: Use when an agent tool (webhook/server, client, system, or MCP) errors, fails validation, won't save, isn't being invoked, returns no result, or behaves wrong, including "tool not used", "property cannot be empty", async/timeout, or test/simulation failures.
---
Diagnose a misbehaving tool by reading the exact error first, then matching the failure pattern. All API calls use `xi-api-key: $API_KEY` against `https://api.elevenlabs.io`.
## Check this FIRST on any schema_mismatch from a webhook or client tool
Every parameter/property object must set EXACTLY ONE value source: a non-empty `description` (the LLM fills it), OR `dynamic_variable`, OR a non-empty `constant_value`, OR `is_system_provided: true`.
- ZERO sources is invalid. `{"type":"string","constant_value":""}` with no description is rejected; an empty-string `constant_value` does NOT count as a constant. Give the field a real `description`.
- TWO OR MORE sources is invalid.
- This applies to `parameters`, `request_body_schema.properties`, `path_params_schema`, `query_params_schema`, and `request_headers` alike.
If the error names `type` or `constant_value` on a property, this rule is almost always the cause. Fix the value source before investigating anything else.
## Do not repeat these mistakes
1. Never guess a tool's config schema or hand-build config JSON blind. Read the tool's actual saved config first with `GET /v1/convai/tools/{tool_id}` and change from there.
2. Read the actual error before theorizing. For a runtime failure, read the conversation where it happened with `GET /v1/convai/conversations/{conversation_id}` and find the failing tool call and its error verbatim; for a save failure, ask the user to quote the exact error text.
3. State tool field names, defaults, and limits from the current config you read, not from memory; they change.
## Match the failure pattern
A. Agent says it uses the tool but never calls it. Check in order: (1) the tool is actually attached to the agent (`GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`, confirm the id is in `tool_ids` or the right node's `additional_tool_ids`); (2) the description is too vague, the model decides invocation from it, so rewrite it to state clearly WHEN to use the tool and apply via `PATCH /v1/convai/tools/{tool_id}`; (3) suggest a simulation test to reproduce deterministically before going live (see the create-simulation-test skill).
B. Validation / won't-save errors ("Property description cannot be empty", missing required field). Read the tool config, find the empty or invalid field, and apply the fix by reading the full config with `GET /v1/convai/tools/{tool_id}` and writing it back with `PATCH /v1/convai/tools/{tool_id}`. For each property confirm the exactly-one-value-source rule above holds.
C. No result returned / async behavior (the tool starts a job but the answer never comes back in the same turn). Likely cause: `execution_mode` is `async` when it should wait, or a client tool's `expects_response` is false. Confirm the current values via `GET /v1/convai/tools/{tool_id}`, then fix with `PATCH /v1/convai/tools/{tool_id}`. See the edit-existing-tool skill for the async plus `expects_response` rule.
D. Webhook tool returns a backend error (a 4xx/5xx, or app-specific text). This is the user's own endpoint failing, NOT an ElevenLabs config bug. Make that distinction early. Confirm from `GET /v1/convai/conversations/{conversation_id}` what payload was sent and what status came back, then verify with the user: endpoint URL, method, auth headers, and that each parameter maps to something the endpoint expects. The fix usually lives in their endpoint or the parameter schema, not the agent.
## When to stop
If the error is from the user's own backend and needs engineering help, or you have looped twice without progress, summarize the exact error and what you have ruled out and hand off to a human (a support form) rather than guessing schemas.
architect-turn-taking-latency1.63 KB
---
name: architect-turn-taking-latency
description: Use when the user asks about turn taking, speculative turns, agent response latency, preparing responses while the user speaks, turn eagerness or timeout, or how fast the agent replies.
---
# Turn taking and latency settings
Everything is branch-scoped over the ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`). You need `$AGENT_ID` and `$BRANCH_ID`.
## 1. Read the current turn settings
Read the agent's `conversation_config.turn` (and the workflow nodes if the question is workflow-specific) with `GET /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID`. Do not pull the full agent context for procedures and unrelated config; the turn block is what matters here.
## 2. Explain the relevant levers
Based on the user's specific latency or turn-taking question:
- Speculative turn (`speculative_turn`): when enabled, the agent starts preparing the LLM response during silence before full turn confidence is reached, reducing perceived latency.
- Turn eagerness (`turn_eagerness`): controls how quickly the agent jumps in (low, standard, high / eager).
- Turn timeout (`turn_timeout`): maximum wait time before the agent re-engages.
## 3. Apply the change
Use a single-field partial-merge patch with the exact dotted path so you touch only the setting in question:
- `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` with `{"conversation_config":{"turn":{"speculative_turn":true}}}`, or the same shape for `turn_eagerness` / `turn_timeout`.
This commits to branch HEAD and returns a new version. Confirm the change to the user, and note it is on the branch until merged.
architect-update-config-safely6.25 KB
---
name: architect-update-config-safely
description: Use when changing an agent setting (LLM, voice, TTS model, language, first message, ASR, guardrails, timeouts, turn-taking, data collection, evaluation criteria) and wanting it to actually apply, or on symptoms like "it won't save", "the change didn't take", "schema mismatch", "invalid config", "config validation failed", "set X to Y", or a language/voice change that broke the audio.
---
# Update agent config safely
Agent config edits fail often, most commonly with a schema mismatch that names no field, so a blind retry just fails again. The reliable path is: read the current state first, change one well-formed field at a time, and know the cross-field constraints before you send. Follow this every time you touch agent config.
Host `https://api.elevenlabs.io`, header `xi-api-key: $API_KEY`. The engineer supplies `$API_KEY`, `$AGENT_ID`, and `$BRANCH_ID`.
## 1. Read the current state first
Fetch the branch-scoped config before writing anything:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" \
-H "xi-api-key: $API_KEY"
```
This gives you the authoritative current config: the exact dotted path you intend to change, the current value you are about to overwrite, and the surrounding settings that constrain the edit (`conversation_config.agent.language`, `conversation_config.tts.model_id`, `conversation_config.agent.prompt.llm`, `voice_id`). The workflow ships in the same response. If the agent has a workflow, config can live partly on nodes (per-node prompt, additional tool ids, additional knowledge base), so confirm whether your change belongs on the base agent or on a node before you patch. Do not edit blind.
Where an edit references tools, read them via `GET /v1/convai/tools` and `GET /v1/convai/tools/{tool_id}`. Where you plan to re-run tests after the change, list them first. Make independent reads in parallel.
Do not invent config keys, enum values, or schema paths. Confirm the exact key name and legal value from the current config response, not from memory.
## 2. Prefer one field per PATCH
A partial-merge `PATCH /v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID` commits straight to branch HEAD and returns a new `version_id`. Send the single leaf you want, nested in its real shape, rather than resending a big config blob. Common partial bodies:
- llm: `{"conversation_config":{"agent":{"llm":"..."}}}`
- tts model: `{"conversation_config":{"tts":{"model_id":"eleven_flash_v2_5"}}}`
- language: `{"conversation_config":{"agent":{"language":"..."}}}`
- first message: `{"conversation_config":{"agent":{"first_message":"..."}}}`
- turn taking: `{"conversation_config":{"turn":{"speculative_turn":true}}}`
- criteria: `{"platform_settings":{"evaluation":{"criteria":[...]}}}` (send the full array; each `conversation_goal_prompt` max 2000 chars)
- data collection: `{"conversation_config":{"platform_settings":{"data_collection":{...}}}}`
Add `&version_description=...` to label the version.
```bash
curl -s -X PATCH "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID&version_description=set-tts-model" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"conversation_config":{"tts":{"model_id":"eleven_flash_v2_5"}}}'
```
For a large text field (system prompt, first message, tool description, node prompt, procedure content), do not push the whole field through a config PATCH. Route to the string-field edit flow instead. See the `architect-edit-string-fields` skill.
Tool bodies (webhook/client/code tools) are not agent-config edits. Update them via `PATCH /v1/convai/tools/{tool_id}` (send the full tool config; it replaces). Workflow nodes and edges are not agent-config edits either.
## 3. Cross-field constraints and gotchas
- Language and TTS model move together (the big one). Changing `conversation_config.agent.language` to a non-English language while `conversation_config.tts.model_id` is an English-only model (`eleven_flash_v2`, `eleven_turbo_v2`) is rejected by validation, or the agent goes silent or garbled at call time. Always pair them: switch `model_id` to a multilingual model (`eleven_flash_v2_5`, `eleven_turbo_v2_5`, or an `eleven_v3` model) in the same change. Send both fields in one PATCH body, or send the model first and the language second.
- Voice vs model: some voices are trained for specific model families. If audio degrades after a model swap, verify `voice_id` is compatible.
- Send each value in its native type: booleans as real booleans, numbers as numbers, enums as the exact allowed string. A number sent as a string is a common silent schema mismatch.
- Enum-valued fields (execution modes, `transfer_type`, ASR quality, LLM model id) must match the allowed set exactly, e.g. `transfer_type` is `blind` or `conference`, never `warm`/`cold`.
- Character limits exist (`conversation_goal_prompt` about 2000 chars, data-collection field descriptions about 500 chars). Over-length values fail validation.
## 4. Recover by error
- `schema_mismatch` with no field detail: the payload shape is wrong. Do not blind-retry. Re-read the exact current shape, then send one field in its real nesting. If it was a language change, apply the language and TTS-model pairing above.
- `validation`: a value is out of range or format (too long, wrong enum, wrong type). Fix that one value against the limits above.
- `not_found`: wrong agent id or branch, a path that does not exist in this config, or a referenced id (tool id, voice id) not on this agent. Re-read the config to get valid paths and ids, and confirm you are on the intended branch.
- transient backend error: retry once. If it persists, treat it as one of the above.
## 5. Verify and follow up
If the change did not seem to take, re-read that single path from the branch-scoped config and compare. A value that reads back unchanged usually means a silent validation rejection or a cross-field constraint, not a glitch.
After a behavioral change, consider re-running the relevant test to confirm no regression. If the change is risky on a live agent, do it on a branch and split traffic rather than editing main directly. For which value to pick for a given setting (LLM choice, voice settings, security posture), defer to the topic-specific skill and use this one only for how to apply the edit safely.
code-tools8.26 KB
---
name: code-tools
description: "Generate JavaScript/TypeScript modules for ElevenLabs code tools — sandbox tools that export a default async (ctx) => function, call APIs via fetch with ctx.secrets / ctx.config / ctx.auth_connections, and return a minimized result to the agent. Use when the user says 'write a code tool', 'code tools', 'ElevenLabs code tool', 'create a code tool', 'egress with auth connection', 'code tool allowed domains', 'X-With-Auth-Connection', or pastes a brief for custom JS/TS tool logic that should run on ElevenLabs infrastructure."
---
# ElevenLabs Code Tools
Generate paste-ready JS/TS for ElevenLabs **code tools**: custom logic that runs on ElevenLabs infrastructure, not a customer webhook.
## Canonical runtime model
Your code is a single JavaScript/TypeScript module that exports one default async function. It receives a `ctx` object and returns the tool result (same downstream uses as a server tool response: agent, transcript, dynamic variables).
```ts
export default async (ctx) => {
const { city } = ctx.args;
return { message: `Hello from ${city}!` };
};
```
Helpers and top-level variables outside the default export are fine.
### `ctx` API (use these names only)
| Property | Description |
|---|---|
| `ctx.args.<paramName>` | Tool-call parameters from the LLM, already typed (string, number, boolean, object). |
| `ctx.secrets.<NAME>` | Workspace secret mapped into this tool's Context object. |
| `ctx.config.<NAME>` | Plain string env/config value mapped into Context. |
| `ctx.auth_connections.<NAME>` | Auth-connection reference. Pass as the `X-With-Auth-Connection` header on outbound `fetch`; ElevenLabs attaches the credential — the raw secret never enters your code. |
Which secrets, auth connections, and config values appear under `ctx` (and under what name) is configured in the tool's **Context object** section in the studio — same idea as headers/params on a webhook tool.
### Naming gotchas (do not emit)
| Wrong | Right |
|---|---|
| `ctx.pargs` | `ctx.args` |
| `ctx.authConnections` | `ctx.auth_connections` |
### Hard constraints
- **No npm packages.** Built-in JavaScript + Web APIs only (`fetch`, `Promise`, `setTimeout`, `JSON`, `URL`, `encodeURIComponent`, etc.). No `import` from external packages.
- **Network allowlist.** Outbound requests only succeed for domains listed under the agent's **General Settings → Code tool allowed domains** (wildcards like `*.example.com` OK). Only workspace **admins** can add domains — flag that if the user is not an admin.
- **Timeout.** Each run must finish within the tool's response timeout (**1–30 seconds**). Cap any poll loops so total wait fits under that budget.
- **Return value = LLM context.** Everything you return is read by the model. Project to the few fields the agent needs — never dump full DB rows, internal notes, risk scores, audit logs, or other PII/internal fields.
- **Auth.** API keys → `ctx.secrets`. Plain URLs/base strings → `ctx.config`. OAuth / managed credentials → `ctx.auth_connections` + `"X-With-Auth-Connection"` header. Never hardcode secrets or put raw OAuth tokens in code.
- **Errors.** Both a structured return and a throw "work" — the difference is control. A structured `{ error: "..." }` return is always visible to the agent; `throw new Error(...)` lets you decide whether the failure reason surfaces to the LLM at all. Prefer structured returns for expected failures (not found, bad input); throw for unexpected upstream failures where a loud transcript failure helps debugging. Either way, try/catch and return or throw something meaningful. Best-effort side effects (logging) should swallow errors so they never break the tool.
## Intake
Ask only for what's missing:
1. **Tool name** + one-line purpose
2. **Parameters** (name, type, description) the LLM will fill → become `ctx.args`
3. **Upstream APIs / domains** to call
4. **Auth style** — secret vs auth connection vs plain config — and suggested Context object names
5. **Return shape** — fields the agent should see
If the brief is already complete, skip questions and generate.
## Output template (always)
Produce all six sections every time:
### 1. Code
Complete default-export module ready to paste into the studio code editor.
### 2. Parameters
Table for Setup params UI:
| Name | Type | Description |
|---|---|---|
| … | string / number / boolean / … | … |
### 3. Context object mapping
| `ctx` path | Kind | Maps to |
|---|---|---|
| `ctx.secrets.EXAMPLE_API_KEY` | secret | workspace secret … |
| `ctx.config.SUPABASE_URL` | value | literal / config string … |
| `ctx.auth_connections.EXAMPLE_CRM` | auth connection | configured connection … |
### 4. Allowed domains checklist
Exact hostnames or wildcards an admin must add under **General Settings → Code tool allowed domains**. Call out that only admins can edit this list.
### 5. Run-in-editor test plan
- Sample **Params** values
- Expected **Output**
- Optional: `console.log(ctx)` / specific fields if debugging
### 6. Agent-facing notes
One short tool description / trigger guidance for when the LLM should call this tool.
## Generation patterns
Keep recipes short here; full sources live in [examples.md](examples.md). Read that file when implementing a matching pattern.
1. **Field-minimizing REST lookup** — fetch a rich row; return only agent-safe fields.
2. **Parallel fan-out** — independent calls via `Promise.all`.
3. **Bounded poll / long-running job** — start work, poll with capped attempts + delay; return `still_*` + id if not done.
4. **Best-effort side-effect via auth connection** — `X-With-Auth-Connection`; try/catch so logging never fails the tool.
5. **Secret-based API key call** — `Authorization` / `apikey` from `ctx.secrets`.
## Anti-patterns (reject or fix)
- Returning entire upstream JSON blobs — return only the fields strictly necessary for the agent; if it's genuinely unclear which fields matter, returning the rest of the object is fine (sometimes that's the intended behavior), but PII/internal fields (notes, risk scores, audit logs) still must never be returned, per the hard constraint above
- Hardcoding secrets or base URLs that belong in Context
- Calling domains without listing them in the Allowed domains checklist
- Polling without attempt/time caps relative to the 30s ceiling
- Putting real OAuth tokens in code instead of `X-With-Auth-Connection`
- Adding npm `import`s
- Using `ctx.pargs` or `ctx.authConnections`
## Minimal egress example (secret)
```ts
export default async (ctx) => {
const orderId = String(ctx.args?.order_id ?? "").trim();
if (!orderId) return { error: "order_id is required" };
const response = await fetch(
`https://api.example.com/orders/${encodeURIComponent(orderId)}`,
{
headers: {
Authorization: `Bearer ${ctx.secrets.EXAMPLE_API_KEY}`,
},
},
);
if (!response.ok) {
throw new Error(`Upstream error: ${response.status}`);
}
const order = await response.json();
return { orderId: order.id, status: order.status };
};
```
Map `EXAMPLE_API_KEY` in Context; allowlist `api.example.com`. Return only the fields the agent needs.
## Minimal OAuth / auth-connection example
```ts
export default async (ctx) => {
const customerId = String(ctx.args?.customer_id ?? "").trim();
if (!customerId) return { error: "customer_id is required" };
const response = await fetch(
`https://api.example.com/customers/${encodeURIComponent(customerId)}`,
{
headers: {
"X-With-Auth-Connection": ctx.auth_connections.EXAMPLE_CRM,
},
},
);
if (!response.ok) {
throw new Error(`Upstream error: ${response.status}`);
}
const customer = await response.json();
return { customerId: customer.id, name: customer.name };
};
```
Map `EXAMPLE_CRM` to a configured auth connection in Context. Return only the fields the agent needs.
## Testing reminder
Before saving in the studio, use **Run** in the code editor:
- **Params** — test values for each defined parameter
- **Output** — returned result or error
- **Logs** — `console.log` / `warn` / `error`, plus build and execution timing
Optional pre-studio check: `node local-executor.mjs <tool-file> <ctx.json>` (see [README.md](README.md#run-locally-before-pasting-into-the-studio)) runs the module locally against a mocked `ctx`, enforcing the same 1–30s timeout.
Referenced files: 4
creative-studio3.28 KB
--- name: creative-studio description: Generate and edit media through the connected ElevenLabs MCP server — speech, images, video, music, and transcription. Use when the user asks to generate a voiceover, image, video, or soundtrack, edit an image, or transcribe audio directly, rather than writing application code. Requires the ElevenLabs MCP server bundled with this plugin. license: MIT --- # ElevenLabs Creative Studio (via MCP) Workflow guidance for the `creative_*` tools on the ElevenLabs MCP server. Every tool takes a `context` parameter — briefly state the user's goal in it. For building media features into the user's own codebase (SDK, API), use the general skills instead (`text-to-speech`, `music`, `sound-effects`, `speech-to-text`). ## How generations work - `creative_generate_speech`, `creative_generate_image`, and `creative_generate_video` return immediately with a `flow_id`, `node_id`, `session_ids`, and a canvas `url` the user can open to keep editing. If the result doesn't render in a view, poll `creative_get_flow_run_status` until `all_completed` or `has_failures` is true. - **Generations spend credits.** Pass `estimate_only: true` to price a long video or a batch of variations before committing. Never call a generation tool a second time to "retry" — that starts and charges a second generation. - `generations_count` defaults to 4 variations so the user can pick; keep the default unless they ask for a specific number. ## Speech - `voice_id` is required and only ever comes from `creative_list_voices` (or the user) — never from memory. Pick a voice matching what the user described; if they gave no hint, pick a clear general-purpose voice rather than asking. - Default model: `eleven_multilingual_v2`. Use `eleven_v3` when the script uses inline audio direction tags like `[whispering]` or `[laughs softly]`. ## Images - Default model: `gemini-2.5-flash-image`. Pick by need: `gpt-image-2` for rendered text, infographics, UI mockups, or reference-driven edits; `flux-2-pro` for fine detail and strict prompt adherence. - `creative_get_flow_node_types` lists what the workspace can run; `creative_get_model_guide` explains how to prompt a specific model. ## Flows: combining generations - Nodes on different flows cannot be connected. When one generation feeds another (lipsync, a voiceover over video), call `creative_create_flow` first and pass that `flow_id` to every related call. - Wire upstream nodes with `connect_from`. A `node_id` only ever comes from a tool result or `creative_get_flow` — never invent one. - `creative_edit_image` requires `connect_from`: a node from an earlier generation on the same flow, a library asset, or an upload. ## Reference files and transcription - For a file already reachable (attached to the conversation, or a direct link): `creative_attach_reference_file` — returns a node with content. - For a file on the user's machine: `creative_upload_flow_reference` — the node stays empty until the user picks a file, so confirm it has an asset (via `creative_get_flow`) before generating from it. - Transcribe with `creative_transcribe_audio`, passing the audio node's id as `connect_from`; without it, a picker handles upload and transcription on its own — don't call the tool again or poll while it's open.
dubbing12.8 KB
---
name: dubbing
description: Dub audio and video into other languages using the ElevenLabs Dubbing API (dubbing_v2), preserving the original speakers' voices. Use when translating videos, podcasts, or recordings into other languages, localizing media content, reviewing or correcting dubbing transcripts and translations, or regenerating a dub after edits.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Dubbing
Dub audio or video into other languages while preserving the original speakers' voices. Create a project from a file or URL, review and edit the source transcript, add one or more target languages, refine translations per segment, and regenerate outputs.
> **Important:** Use the Dubbing Projects API — `elevenlabs.dubbing.project.*` in the SDKs, or the `/v1/dubbing/project` REST endpoints. Do **not** use the legacy v1 dubbing surface (`client.dubbing.create()`, `client.dubbing.get()`, `client.dubbing.audio.get()`, or bare `/v1/dubbing` routes) — that is the older dubbing API, now under Legacy in the API reference.
> **Setup:** See [Installation Guide](references/installation.md). The `elevenlabs` CLI and the SDKs read `ELEVENLABS_API_KEY` automatically; REST base URL is `https://api.elevenlabs.io` with your API key in the `xi-api-key` header.
## Concepts
| Concept | Meaning |
|---------|---------|
| **Project** | One source of media (file or URL) plus its source transcript. Prepared (transcribed) once, then rests in `ready` while you add languages. |
| **Source transcript** | Editable segments (text, speaker, timing) transcribed from the source. The single source of truth every language is translated from. |
| **Language (target)** | One dubbed output language. Each has its own transcript (source segments + a translation per segment) and its own dubbed audio output. |
| **Revisions** | Independent monotonic counters. The project's `revision` bumps on source-transcript edits; a language's `revision` bumps on translation edits or source edits that affect it. A language's `output_revision` is the revision its current audio was generated from — when it's behind `revision`, the output is out of date. |
**Recommended order of operations:** finalize the source transcript **before** adding any languages. Translations are produced from the source, so correcting the source first means every language starts from the right text — editing the source after a language completes marks it `stale` and requires a (charged) regeneration.
> **Enterprise:** Transcript editing and regeneration are available to enterprise workspaces only. Creating projects, adding languages, and downloading dubs work on all plans.
## Workflow
1. **Create** the project from a file or URL → `queued`
2. **Poll** the project until `ready`
3. **Review and finalize the source transcript** (edit/add/delete segments)
4. **Add** one language per target → `queued` → `processing` → `completed`
5. **Download** each language's `outputs.lossless_audio` when `completed`
6. **Refine** translations per segment if needed → the language goes `stale`
7. **Regenerate** the language → `completed` again with fresh output
## Quick Start (Python)
```python
import os
import time
import requests
from elevenlabs.client import ElevenLabs
elevenlabs = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
# 1. Create a project from a local file (or pass source_url=... instead of file)
with open("promo.mp4", "rb") as f:
project = elevenlabs.dubbing.project.create(
file=f,
source_language="en",
reference="Q3 marketing video",
)
# 2. Wait for the source media to be transcribed
while True:
project = elevenlabs.dubbing.project.get(project.project_id)
if project.status == "ready":
break
if project.status == "failed":
raise RuntimeError("Project preparation failed")
time.sleep(5)
# 3. Add a Spanish language target
language = elevenlabs.dubbing.project.language.create(
project.project_id,
target_language="es",
)
# 4. Wait for the dub to finish generating
while True:
language = elevenlabs.dubbing.project.language.get(
project.project_id, language.language_id
)
if language.status == "completed":
break
if language.status == "failed":
raise RuntimeError("Dub generation failed")
time.sleep(5)
# 5. Download the dubbed audio (signed URL, valid ~1 hour — re-fetch the language for a fresh one)
audio = requests.get(language.outputs.lossless_audio)
with open("promo_es.wav", "wb") as f:
f.write(audio.content)
```
## Quick Start (JavaScript)
```typescript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { writeFile } from "fs/promises";
const elevenlabs = new ElevenLabsClient();
// 1. Create a project (sourceUrl shown; file upload is also supported)
let project = await elevenlabs.dubbing.project.create({
sourceUrl: "https://example.com/promo.mp4",
sourceLanguage: "en",
reference: "Q3 marketing video",
});
// 2. Wait for the source media to be transcribed
while (true) {
project = await elevenlabs.dubbing.project.get(project.projectId);
if (project.status === "ready") break;
if (project.status === "failed") throw new Error("Project preparation failed");
await new Promise((resolve) => setTimeout(resolve, 5000));
}
// 3. Add a Spanish language target
let language = await elevenlabs.dubbing.project.language.create(project.projectId, {
targetLanguage: "es",
});
// 4. Wait for the dub to finish generating
while (true) {
language = await elevenlabs.dubbing.project.language.get(project.projectId, language.languageId);
if (language.status === "completed") break;
if (language.status === "failed") throw new Error("Dub generation failed");
await new Promise((resolve) => setTimeout(resolve, 5000));
}
// 5. Download the dubbed audio from the signed URL
const response = await fetch(language.outputs!.losslessAudio!);
await writeFile("promo_es.wav", Buffer.from(await response.arrayBuffer()));
```
## Quick Start (CLI)
The `elevenlabs` CLI reads `ELEVENLABS_API_KEY` from the environment automatically.
```bash
# 1. Create a project (use --source-url "https://..." instead of --file to dub from a URL)
elevenlabs dubbing project create --file promo.mp4 --source-language en
# → {"project_id": "proj_...", "status": "queued", ...}
# 2. Poll until status is "ready"
elevenlabs dubbing project get --project-id proj_...
# 3. Add a target language
elevenlabs dubbing project language create --project-id proj_... --target-language es
# 4. Poll the language until "completed", then download outputs.lossless_audio
elevenlabs dubbing project language get --project-id proj_... --language-id lang_...
```
## Create Options
`elevenlabs dubbing project create` (REST: `POST /v1/dubbing/project`, `multipart/form-data`) takes **either** `file` **or** `source_url` (not both):
| Field | Required | Notes |
|-------|----------|-------|
| `file` | one of file/source_url | Source media to dub (audio or video), up to 3 GiB |
| `source_url` | one of file/source_url | Public URL to fetch the source media from |
| `source_language` | no | ISO 639 code (e.g. `en`). Omit to auto-detect — the detected language is reported on the source transcript's `language` field |
| `reference` | no | Free-form label to identify the project on your end (max 500 chars) |
| `model_id` | no | `dubbing_v2` (default) |
| `target_language` | no | Optionally queue the first language target at creation; add more with `language.create` |
| `keyterms` | no | Terms to bias transcription/translation toward (product/brand names). Up to 1000 terms; each at most 50 chars and 5 words; `<>{}[]\` not allowed. Repeat the field once per term in multipart |
## Editing the Source Transcript
Once the project is `ready`, read the transcript, then correct it before adding languages. Every edit bumps the project's `revision`. Each segment has a stable `id` used to edit or delete it. (Enterprise workspaces only.)
```python
# Read the source transcript
transcript = elevenlabs.dubbing.project.transcript.get(project_id)
# Correct a segment's text — send only the fields to change (text, speaker_id, start_s, end_s)
elevenlabs.dubbing.project.transcript.update_segment(
project_id,
segment_id=transcript.segments[0].id,
text="Welcome to our latest product demo.",
)
# Add a segment (reuse an existing speaker_id so it's dubbed with that speaker's voice)
added = elevenlabs.dubbing.project.transcript.create_segment(
project_id,
text="Thanks for watching.",
speaker_id=transcript.segments[0].speaker_id,
start_s=40.0,
end_s=42.0,
)
# Delete a segment
elevenlabs.dubbing.project.transcript.delete_segment(project_id, segment_id=added.segment.id)
```
Via the CLI: `elevenlabs dubbing project transcript get --project-id proj_...`, then update a segment with only the changed fields (`--text`, `--speaker-id`, `--start-s`, `--end-s`):
```bash
elevenlabs dubbing project transcript update_segment \
--project-id proj_... --segment-id seg_... \
--text "Welcome to our latest product demo."
```
## Refining Translations and Regenerating
A language's transcript pairs each source segment with its `translation` (`null` = not yet translated; segment ids match the source). Edit a single translation, then regenerate. (Enterprise workspaces only.)
```python
# Read the language's translations
target = elevenlabs.dubbing.project.language.transcript.get(project_id, language_id)
# Refine a single translation (pass translation=None to clear it and mark for re-translation)
elevenlabs.dubbing.project.language.transcript.update_segment(
project_id,
language_id,
segment_id=target.segments[0].id,
translation="Bienvenido a nuestra última demostración de producto.",
)
# Regenerate the dub from the current transcript (charged like a generation)
elevenlabs.dubbing.project.language.transcript.regenerate(project_id, language_id)
```
Via the CLI: `elevenlabs dubbing project language transcript update_segment --project-id proj_... --language-id lang_... --segment-id seg_... --translation "..."`, then `elevenlabs dubbing project language transcript regenerate --project-id proj_... --language-id lang_...` (returns `202 Accepted`).
A translation edit affects only that language. After the edit, a `completed` language becomes `stale` — it keeps serving its previous output until you regenerate. Poll until `completed`; `output_revision` then equals `revision` and `outputs.lossless_audio` reflects the current transcript.
## Dubbing into Multiple Languages
Add one language target per language — each generates independently. Track them all with `language.list` instead of polling one by one:
```python
for lang in ["es", "fr", "de", "ja"]:
elevenlabs.dubbing.project.language.create(project_id, target_language=lang)
while True:
result = elevenlabs.dubbing.project.language.list(project_id)
if not any(l.status in ("queued", "processing") for l in result.languages):
break
time.sleep(5)
```
## States
**Project:**
| Status | Meaning |
|--------|---------|
| `queued` | Created; source fetch + preparation enqueued |
| `preparing` | Preparation (transcription) running |
| `ready` | Source transcript available; add/generate languages. Projects **stay** `ready` — per-language progress lives on the languages |
| `failed` | Preparation failed (e.g. source couldn't be fetched or decoded) |
**Language:**
| Status | Meaning |
|--------|---------|
| `queued` | Waiting on the project becoming `ready`, or on a generation worker |
| `processing` | The dub is being generated |
| `completed` | Finished; `outputs` populated with a signed download URL (valid ~1 hour — re-fetch for a fresh one) |
| `stale` | Previously completed, but the transcript changed; keeps the last output until regenerated |
| `failed` | Generation failed |
You can add a language before the project is `ready` — it stays `queued` and starts automatically once the project becomes `ready`. Adding a language accepts optional `model_id` (defaults to the project's) and `voice_settings` (e.g. `{"cloning_strength": 7}`, range 0–10, default 7 — controls how strongly dubbed speakers clone the source voices).
## Error Handling
- **401**: Invalid API key
- **409 Conflict** on regenerate: The project isn't `ready` or the language isn't settled (e.g. already generating) — wait and retry
- **Expired download URL**: `outputs.lossless_audio` is signed and valid ~1 hour; re-fetch the language for a fresh URL
- **Transcript editing / regeneration unavailable**: These endpoints are enterprise-only — on other plans, create the project with a finalized source and add languages directly
## References
- [Installation Guide](references/installation.md)
- [API Reference](references/api-reference.md) — every endpoint with full request/response schemas and SDK method names
Referenced files: 2
fix-agent-qa-ticket13.4 KB
---
name: fix-agent-qa-ticket
description: >
Fix an Architect QA triage ticket (agtqa_*) end-to-end via the raw agents
API: fetch the ticket + conversation, find the root-cause procedure or prompt
text, fix it on a branch, add + run a simulation test proving the fix, then post
a summary comment on the ticket. Use when asked to "fix agtqa_...", "address this
QA ticket", "triage and fix this Architect ticket", or given an agtqa_* id/URL.
---
# Fix a QA ticket (agtqa_*)
`agtqa_*` = an Agent Conversation Triage Ticket: a reviewer's flagged comments
against specific turns of a real Architect/ConvAI conversation. This skill turns
one of those tickets into a verified, branch-scoped fix, entirely through the raw
ConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key`). It never
touches main directly — it stages everything on a branch and leaves merging to
the user.
**Needs from the user**: an API key and a ticket id (`agtqa_...`) or a URL
containing one. Everything else — agent id, conversation id, branch — is derived
from the ticket.
**Platform-only.** Do not write repo code as part of this skill unless the root
cause is actually a code bug (rare — most triage findings are prompt/procedure
gaps). If code IS the root cause, say so and hand off instead of guessing at a
platform-side workaround.
## 1. Fetch the ticket
```bash
curl -s "https://api.elevenlabs.io/v1/convai/conversation-triage-tickets/$TICKET_ID" \
-H "xi-api-key: $API_KEY"
```
Returns `conversation_id`, `agent_id`, `qa_comment`, `ticket_comments[]`,
`turn_comments[]` (each with `turn_index` + `comment`), `status`
(`open`/`in_progress`/`resolved`).
## 2. Fetch the conversation and read the flagged turns
```bash
curl -s "https://api.elevenlabs.io/v1/convai/conversations/$CONVERSATION_ID" \
-H "xi-api-key: $API_KEY"
```
`transcript[]` has per-turn `role`, `message`, `tool_calls` (each with
`tool_name` + `params_as_json`). Read a window around every `turn_index` named in
`turn_comments` — usually a couple turns before and after — to understand what
the agent actually did and said. The reviewer's comments tell you *what's wrong*;
the transcript tells you *why*. Don't skip straight to guessing the fix — find
the specific tool call, prompt gap, or missing instruction that produced the bad
turn.
The conversation also carries `branch_id` and `version_id` — useful context, but
**do not build your fix branch off this conversation's branch** unless it's
demonstrably the right base (check `is_archived` / whether it's actually merged
into main — see step 3). It's just where the flagged conversation happened to run.
## 3. Create a fix branch off main's actual tip
Get the agent and find its real main branch:
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID" -H "xi-api-key: $API_KEY"
# -> .main_branch_id, .branch_id (top-level is usually main)
```
Then fetch that branch to get its true HEAD version (do not assume — a stray
personal/archived branch can look tempting but have unrelated commits in its
ancestry):
```bash
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$MAIN_BRANCH_ID" \
-H "xi-api-key: $API_KEY"
# -> .most_recent_versions[0].id is main's real tip version_id
```
Create the branch from that exact version:
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"name": "<you>/qa-<ticket-suffix>-<short-desc>", "description": "Fix for '"$TICKET_ID"'", "parent_version_id": "<main-tip-version-id>"}'
```
Required fields are `name`, `description`, `parent_version_id` (all three, or
you get a 422 listing what's missing). Response: `{created_branch_id,
created_version_id}`.
## 4. Find and fix the root cause
Usually one of:
- **A procedure** missing an instruction (list via
`GET .../agents/{id}/branches/{b}/procedures`, fetch full content via
`GET .../procedures/{pid}`). Search procedure names/triggers for the relevant
topic (dashboards, refunds, whatever the ticket concerns).
- **The system prompt** (`conversation_config.agent.prompt.prompt` off
`GET /v1/convai/agents/{id}?branch_id={b}`).
### Editing an existing procedure (draft + publish dance)
Editing an existing procedure is **two steps**, not one — there is no direct
commit-on-PATCH for procedures:
```bash
# 1. Stage the draft (full content: frontmatter + body, or per-field if the
# procedure isn't markdown-frontmatter style)
curl -s -X PATCH \
"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID/draft" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"name": "...", "content": "...", "type": "free_form", "trigger": "..."}'
# 2. Publish ALL pending procedure drafts on the branch by re-committing the
# agent with any partial merge body (re-sending the branch's current,
# unmodified prompt is the simplest no-op payload). Fetch the CURRENT prompt
# scoped to $BRANCH_ID first (not main's) so you don't clobber other
# branch-local differences:
curl -s "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID" -H "xi-api-key: $API_KEY" \
| jq -r '.conversation_config.agent.prompt.prompt' > current_prompt.txt
curl -s -X PATCH "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID&version_description=..." \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d "{\"conversation_config\": {\"agent\": {\"prompt\": {\"prompt\": $(jq -Rs . < current_prompt.txt)}}}}"
```
Verify by re-`GET`ting the procedure and confirming the `version_id` changed and
the new content is present.
A direct `PATCH .../procedures/{pid}` (no `/draft` suffix) does not exist —
returns 405. Creating a brand-new procedure via `POST .../procedures` commits
immediately, but don't use that to "replace" an existing one in place — it
leaves two procedures with overlapping/duplicate triggers, which is worse than
the bug you're fixing.
### Editing the prompt / criteria directly
No draft dance needed — a partial-merge `PATCH /v1/convai/agents/{id}?branch_id={b}`
commits straight to branch HEAD and returns a new `version_id`:
- prompt: `{"conversation_config":{"agent":{"prompt":{"prompt":"..."}}}}`
- criteria: `{"platform_settings":{"evaluation":{"criteria":[...]}}}` (send the
full array; each `conversation_goal_prompt` max 2000 chars)
## 5. Add a simulation test that reproduces the ticket's scenario
**Always use a `simulation` test for Architect/DOM agents** — `llm`/`response`
tests only grade the agent's first action, which for Architect is almost always
`start_procedure`, so the judge returns useless `unknown`/`failure` verdicts.
```bash
# Optional: group tests for this ticket
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agent-testing/folders" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"name": "QA '"$TICKET_ID"'"}'
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agent-testing/create" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{
"type": "simulation",
"name": "<short description> ('"$TICKET_ID"')",
"parent_folder_id": "<folder id or omit>",
"simulation_scenario": "<persona + exactly what they do, mirroring the ticket>",
"success_conditions": ["<checklist item 1>", "<checklist item 2>", ...],
"simulation_max_turns": 8,
"tool_mock_config": {"mocking_strategy": "all", "fallback_strategy": "raise_error"},
"chat_history": [{"role": "user", "message": "...", "time_in_call_secs": 0}],
"dynamic_variables": {"tier": "enterprise", "product": "conversational_ai", "objective": "...", "userInfo": "{}", "chatHistory": "[]"}
}'
```
Gotchas that produce 422s:
- Every `chat_history` entry needs `time_in_call_secs` (even `0`), or you get a
`missing` error pointing at that field.
- `chat_history` must end with a user message.
- A `system`-type `tool_result` inside `chat_history` needs `is_error` +
`tool_has_been_called` fields.
- `dynamic_variables` should include whatever the prompt templates on
(tier/product/objective/userInfo/chatHistory at minimum) or templating crashes
mid-run.
- `mocking_strategy` is an enum of exactly `all` / `selected` / `none` — there
is no fourth value. **`"all"` ignores `tool_mock_overrides` entirely** — it
mocks every tool call with the generic fallback error, even if you've set a
per-tool `mock_result`. If the fix you're testing depends on a specific tool
actually returning realistic data (e.g. `list_branches` returning a branch
list so the agent can resolve a name to an id), you MUST use
`mocking_strategy: "selected"` with `mocked_tool_ids: [...]` naming every
tool you want mocked — `"all"` silently no-ops your overrides and you'll see
"no mock matched" in the transcript even though the override is saved
correctly on the test (verify via `GET /v1/convai/agent-testing/{test_id}`
if a mock isn't taking effect — the stored config can look right while the
run still fails, which means the strategy, not the override, is wrong).
- `tool_mock_overrides` shape: `{"<tool_name>": [{"mock_result":
"<json-string>", "parameter_conditions": [], "is_error": false}]}`.
`mock_result` is a JSON-encoded **string**, not a nested object — build it
with `json.dumps(...)` (or equivalent) before embedding it in the request
body. An empty `parameter_conditions` list means "match unconditionally."
- `mocking_strategy: "all"` + `fallback_strategy: "raise_error"` (no overrides)
makes *every* tool call error mid-conversation ("technical difficulties").
Fine only for testing something that doesn't depend on a tool succeeding
(e.g. a confirmation-message wording fix where the tool call itself is
incidental); if the flow needs a tool to actually succeed with realistic
data, use `selected` + overrides instead of reaching for `call_real_tool`.
- DOM-write tools (navigate, set_dashboard_filters, etc.) still may not fully
green in the sim harness even with `selected` mocking if you haven't mocked
every tool in the chain — treat persistent failures there as behavior
documentation, not a bug, once you've confirmed the mocking strategy itself
isn't the culprit.
Write the `success_conditions` directly against the reviewer's `turn_comments` —
each comment should map to a checklist item the grader can verify. Word each
condition to name the CORRECT id/value explicitly (e.g. "uses id X, not the
raw name string and not id Y") rather than just "doesn't guess" — a vague
condition lets a new, differently-wrong failure mode (e.g. passing the name
itself as the id) slip through as a pass.
**Editing a sim test = delete + recreate.** There is no in-place edit for a
simulation test (`PUT` on a test id is response/llm-only and rejects
`type: simulation`). If a run reveals your mock config was wrong, `DELETE
/v1/convai/agent-testing/{test_id}`, fix the payload, and `POST .../create`
again — then re-attach and re-run.
Attach it to the agent/branch, then run it:
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/testing/attach-test" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"test_id": "'"$TEST_ID"'", "branch_id": "'"$BRANCH_ID"'"}'
curl -s -X POST "https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"tests": [{"test_id": "'"$TEST_ID"'"}], "branch_id": "'"$BRANCH_ID"'", "repeat_count": 1}'
# -> {id: suite_id, test_runs: [{test_run_id, status: "pending", ...}]}
```
## 6. Poll for the result
Sim runs take a few minutes. Poll `GET
/v1/convai/test-invocations/{suite_id}` until `test_runs[0].status` leaves
`pending`. Use a background poll loop (Bash `run_in_background` or Monitor), not
a foreground sleep — do not busy-wait in the conversation.
Once terminal, read `test_runs[0].condition_result.result`
(`success`/`failure`) and `.rationale.messages` (per-criterion grader
reasoning) to confirm the fix actually produces the intended behavior — don't
just check the pass/fail bit, skim the rationale for whether it's testing what
you think it's testing.
If it fails: re-read the rationale against the actual procedure/prompt change,
adjust, and re-run. Don't loosen `success_conditions` to force a pass unless
the condition itself was wrong (too strict/loose) — a passed test the reviewer's
concern doesn't check is worse than an honest failure.
## 7. Comment on the ticket
```bash
curl -s -X POST "https://api.elevenlabs.io/v1/convai/conversation-triage-tickets/$TICKET_ID/comments" \
-H "xi-api-key: $API_KEY" -H "Content-Type: application/json" \
-d '{"comment": "<root cause>\n\n<what changed, branch id, not merged>\n\n<test id + PASS/FAIL + key rationale line>\n\nWritten by <model>, using Claude Code."}'
```
Include: the root cause in plain language, the branch id (explicitly "not
merged" — merging is the user's call), the test id and result, and identify
yourself per repo convention (`Written by {Model}, using {Harness}`).
**Leave ticket `status` as `open`** (don't PATCH it to `resolved`) unless the
user explicitly asks you to close it — you fixed and verified on a branch, but
the user still needs to review and merge.
## 8. Report back
Tell the user: root cause, branch id + that it's unmerged, test id + result,
and anything you had to work around (wrong-parent branch mistake, a blocked
destructive action, an ambiguous base branch) — these are exactly the kind of
judgment calls the user should sanity-check before merging.
music13.2 KB
---
name: music
description: Generate music using ElevenLabs Music API. Use when creating instrumental tracks, songs with lyrics, background music, jingles, or any AI-generated music composition. Supports prompt-based generation, composition plans for granular control, and detailed output with metadata.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Music Generation
Generate music from text prompts - supports instrumental tracks, songs with lyrics, and fine-grained control via composition plans.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
All examples below default to `music_v2`, the current generation model. Pass `model_id="music_v1"` only when explicitly requested to.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio = client.music.compose(
prompt="A chill lo-fi hip hop beat with jazzy piano chords",
music_length_ms=30000,
model_id="music_v2",
)
with open("output.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
```
### TypeScript
```typescript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
const client = new ElevenLabsClient();
const audio = await client.music.compose({
prompt: "A chill lo-fi hip hop beat with jazzy piano chords",
musicLengthMs: 30000,
modelId: "music_v2",
});
audio.pipe(createWriteStream("output.mp3"));
```
### CLI
```bash
elevenlabs music compose \
--prompt "A chill lo-fi beat" \
--music-length-ms 30000 \
--model-id music_v2 \
--output output.mp3
```
## Methods
| Method | Description |
|--------|-------------|
| `music.compose` | Generate audio from a prompt or composition plan |
| `music.stream` | Stream audio chunks as they are generated (paid plans) |
| `music.composition_plan.create` | Generate a structured plan for fine-grained control |
| `music.compose_detailed` | Generate audio + composition plan + metadata; pass `store_for_inpainting=True` to enable inpainting |
| `music.compose_detailed_stream` | Stream audio plus composition plan, metadata, and optional word timestamps as Server-Sent Events |
| `music.video_to_music` | Generate background music from one or more uploaded video files |
| `music.upload` | Upload an audio file for later inpainting workflows, optionally extracting its composition plan or word-level timestamps |
| `music.finetunes.list` | List accessible music finetunes |
| `music.finetunes.create` | Train a music finetune from uploaded audio |
| `music.finetunes.get` | Retrieve finetune status and metadata |
| `music.finetunes.update` | Update finetune metadata or visibility |
| `music.finetunes.delete` | Delete a music finetune |
See [API Reference](references/api_reference.md) for full parameter details.
`music.upload` is available to enterprise clients with access to the inpainting feature.
## Music Finetunes
Create a finetune from training audio with
[`POST /v1/music/finetunes`](https://elevenlabs.io/docs/api-reference/music/finetunes/create),
then poll the [get endpoint](https://elevenlabs.io/docs/api-reference/music/finetunes/get) until
its status is `completed`. Pass the returned `id` as `finetune_id` when composing music.
Use the [list](https://elevenlabs.io/docs/api-reference/music/finetunes/list),
[update](https://elevenlabs.io/docs/api-reference/music/finetunes/update), and
[delete](https://elevenlabs.io/docs/api-reference/music/finetunes/delete) endpoints to manage
accessible finetunes.
## Video to Music
Generate background music from uploaded video clips via
[`POST /v1/music/video-to-music`](https://elevenlabs.io/docs/api-reference/music/video-to-music)
(`client.music.video_to_music`). This is separate from prompt-based
[`music.compose`](https://elevenlabs.io/docs/api-reference/music/compose) (`POST /v1/music`).
The API combines videos in order, accepts an optional natural-language description, and lets you
steer style with up to 10 tags such as `upbeat` or `cinematic`. This endpoint still defaults to
`music_v1`; pass `model_id="music_v2"` to use the newer model.
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio = client.music.video_to_music(
videos=["trailer.mp4"],
description="Build suspense, then resolve with a warm cinematic finish.",
tags=["cinematic", "suspenseful", "uplifting"],
model_id="music_v2",
)
with open("video-score.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
```
### TypeScript
```typescript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";
const client = new ElevenLabsClient();
const audio = await client.music.videoToMusic({
videos: [createReadStream("trailer.mp4")],
description: "Build suspense, then resolve with a warm cinematic finish.",
tags: ["cinematic", "suspenseful", "uplifting"],
modelId: "music_v2",
});
audio.pipe(createWriteStream("video-score.mp3"));
```
### CLI
```bash
elevenlabs music video_to_music \
--videos trailer.mp4 \
--description "Build suspense, then resolve with a warm cinematic finish." \
--tags cinematic \
--model-id music_v2 \
--output video-score.mp3
```
The CLI currently accepts one `--videos` file and one `--tags` value per request; use the Python
or TypeScript SDK to send multiple videos or tags.
Constraints from the current API schema:
- Upload 1-10 video files per request
- Keep total combined upload size at or below 200 MB
- Keep total combined video duration at or below 600 seconds
- Use `description` for high-level musical direction and `tags` for concise style cues
## Composition Plans
`music_v2` composition plans are an ordered list of `chunks`. Each chunk specifies its own
`text` (section label, lyrics, inline cues), `duration_ms`, `positive_styles`, `negative_styles`,
and `context_adherence` (`low`, `medium`, or `high`, default `high`). Up to 30 chunks per plan,
each 3,000–120,000 ms, total length 3 s to 10 minutes.
Generate a plan first, edit it, then compose:
```python
plan = client.music.composition_plan.create(
prompt="An epic orchestral piece building to a climax",
music_length_ms=60000,
model_id="music_v2",
)
# Edit chunks in place
plan["chunks"][0]["text"] = "[Intro]\nQuiet strings rising"
audio = client.music.compose(
composition_plan=plan,
model_id="music_v2",
)
```
```typescript
const plan = await client.music.compositionPlan.create({
prompt: "An epic orchestral piece building to a climax",
musicLengthMs: 60000,
modelId: "music_v2",
});
plan.chunks[0].text = "[Intro]\nQuiet strings rising";
const audio = await client.music.compose({
compositionPlan: plan,
modelId: "music_v2",
});
```
Or hand-build a plan to control lyrics and style per section:
```python
composition_plan = {
"chunks": [
{
"text": "[Verse]\nWalking down an empty street",
"duration_ms": 15000,
"positive_styles": ["pop", "upbeat", "female vocals", "acoustic guitar"],
"negative_styles": ["dark", "slow"],
"context_adherence": "high",
},
{
"text": "[Chorus]\nThis is my moment",
"duration_ms": 15000,
"positive_styles": ["powerful vocals", "full band"],
"negative_styles": [],
"context_adherence": "high",
},
]
}
audio = client.music.compose(composition_plan=composition_plan, model_id="music_v2")
```
```typescript
const compositionPlan = {
chunks: [
{
text: "[Verse]\nWalking down an empty street",
durationMs: 15000,
positiveStyles: ["pop", "upbeat", "female vocals", "acoustic guitar"],
negativeStyles: ["dark", "slow"],
contextAdherence: "high",
},
{
text: "[Chorus]\nThis is my moment",
durationMs: 15000,
positiveStyles: ["powerful vocals", "full band"],
negativeStyles: [],
contextAdherence: "high",
},
],
};
const audio = await client.music.compose({
compositionPlan,
modelId: "music_v2",
});
```
Put broader characteristics (genre, instrumentation, vocal style) in `positive_styles`, not in
`text`. The first chunk's styles set the overall tone — include 6–7 styles there.
## Output Formats
Use the `output_format` query parameter on compose, detailed compose, or stream requests to select
the generated audio format. `auto` chooses a model-appropriate MP3 format; for `music_v2`, it
selects `mp3_48000_192`. Higher-bitrate MP3 options include `mp3_48000_240` and `mp3_48000_320`.
## Streaming
For paid plans, stream audio chunks as they are generated instead of waiting for the full file:
```python
from io import BytesIO
stream = client.music.stream(
prompt="A driving synthwave track with arpeggiated leads",
music_length_ms=30000,
model_id="music_v2",
)
buffer = BytesIO()
for chunk in stream:
if chunk:
buffer.write(chunk)
```
```typescript
const stream = await client.music.stream({
prompt: "A driving synthwave track with arpeggiated leads",
musicLengthMs: 30000,
modelId: "music_v2",
});
const chunks: Buffer[] = [];
for await (const chunk of stream) {
chunks.push(chunk);
}
```
### Detailed streaming
Use detailed streaming when the application needs generated music metadata while audio is still
arriving. `POST /v1/music/detailed/stream` accepts the same prompt or composition-plan body as
detailed compose, streams `text/event-stream`, and can include word timestamps with
`with_timestamps`.
```bash
elevenlabs music compose_detailed_stream \
--prompt "A bright indie pop hook with warm guitars" \
--music-length-ms 30000 \
--model-id music_v2 \
--with-timestamps true \
--output-format auto
```
## Inpainting
Inpainting edits or extends a stored song by mixing **audio reference chunks** (unchanged slices
of a stored song) with new **generation chunks** in a single composition plan.
Step 1 — get a `song_id`, either by storing a fresh generation or uploading existing audio:
```python
# Option A: keep a generation for later editing
result = client.music.compose_detailed(
prompt="An upbeat pop song with verse and chorus",
music_length_ms=60000,
model_id="music_v2",
store_for_inpainting=True,
)
song_id = result.song_id
# Option B: upload an existing track and extract its plan
uploaded = client.music.upload(
file=open("my-song.mp3", "rb"),
extract_composition_plan="music_v2",
)
song_id = uploaded.song_id
composition_plan = uploaded.composition_plan
```
```typescript
import { createReadStream } from "fs";
// Option A: keep a generation for later editing
const result = await client.music.composeDetailed({
prompt: "An upbeat pop song with verse and chorus",
musicLengthMs: 60000,
modelId: "music_v2",
storeForInpainting: true,
});
let songId = result.songId;
// Option B: upload an existing track and extract its plan
const uploaded = await client.music.upload({
file: createReadStream("my-song.mp3"),
extractCompositionPlan: "music_v2",
});
songId = uploaded.songId;
const compositionPlan = uploaded.compositionPlan;
```
Step 2 — compose a plan that references the stored audio and regenerates the part you want to
change:
```python
plan = {
"chunks": [
{"song_id": song_id, "range": {"start_ms": 0, "end_ms": 30000}},
{
"text": "[Chorus]\nWe're rising up tonight",
"duration_ms": 30000,
"positive_styles": ["bigger drums", "layered vocals", "anthemic"],
"negative_styles": ["sparse"],
"context_adherence": "high",
},
]
}
audio = client.music.compose(composition_plan=plan, model_id="music_v2")
```
```typescript
const plan = {
chunks: [
{ songId, range: { startMs: 0, endMs: 30000 } },
{
text: "[Chorus]\nWe're rising up tonight",
durationMs: 30000,
positiveStyles: ["bigger drums", "layered vocals", "anthemic"],
negativeStyles: ["sparse"],
contextAdherence: "high",
},
],
};
const audio = await client.music.compose({
compositionPlan: plan,
modelId: "music_v2",
});
```
To match the feel of a stored slice without copying it, attach a `conditioning_ref` (up to
30,000 ms) plus a `condition_strength` of `low`, `medium`, `high`, or `xhigh` to a generation
chunk. Conditioning placed on the first chunk influences every later chunk.
See [API Reference](references/api_reference.md) for the full inpainting parameter list.
## Content Restrictions
- Cannot reference specific artists, bands, or copyrighted lyrics
- `bad_prompt` errors include a `prompt_suggestion` with alternative phrasing
- `bad_composition_plan` errors include a `composition_plan_suggestion`
## Error Handling
```python
try:
audio = client.music.compose(prompt="...", music_length_ms=30000)
except Exception as e:
print(f"API error: {e}")
```
```typescript
try {
const audio = await client.music.compose({
prompt: "...",
musicLengthMs: 30000,
});
} catch (err) {
console.error("API error:", err);
}
```
Common errors: 401 (invalid key), 422 (invalid params), 429 (rate limit).
## References
- [Installation Guide](references/installation.md)
- [API Reference](references/api_reference.md)
Referenced files: 2
setup-api-key3.76 KB
--- name: setup-api-key description: Guides users through setting up an ElevenLabs API key for ElevenLabs MCP tools. Use when the user needs to configure an ElevenLabs API key, when ElevenLabs tools fail due to missing API key, or when the user mentions needing access to ElevenLabs. First checks whether ELEVENLABS_API_KEY is already configured and valid, and only runs full setup when needed. license: MIT compatibility: Requires internet access to elevenlabs.io and api.elevenlabs.io. --- # ElevenLabs API Key Setup Guide the user through obtaining and configuring an ElevenLabs API key. ## Workflow ### Step 0: Check for an existing API key first Before asking the user for a key, check for an existing `ELEVENLABS_API_KEY`: 1. Check whether `ELEVENLABS_API_KEY` exists in the current environment. If it does, use that value for this initial check. 2. Only if it is not in the environment, check `.env` for `ELEVENLABS_API_KEY=<value>`. 3. Do not print, quote, or repeat the key. If you mention it, redact it. 4. If an existing key is found, validate it: ```text GET https://api.elevenlabs.io/v1/user Header: xi-api-key: <existing-api-key> ``` 5. **If existing key validation succeeds:** - Tell the user ElevenLabs is already configured and working - Skip the setup flow - Ask whether they want to replace/rotate the key; if not, stop 6. **If existing key validation fails:** - Tell the user the existing key appears invalid or expired - Continue to Step 1 ### Step 1: Request the API key Tell the user: > To set up ElevenLabs, open the API keys page: https://elevenlabs.io/app/settings/api-keys > > (Need an account? Create one at https://elevenlabs.io/app/sign-up first) > > If you don't have an API key yet: > 1. Click "Create key" > 2. Name it (or use the default) > 3. Set permission for your key. If you provide a key with "User" permission set to "Read" this skill will automatically verify if your key works > 4. Click "Create key" to confirm > 5. **Copy the key immediately** - it's only shown once! > > Do not paste the key into this chat. Instead, copy/paste it into your local `.env` file: > > ``` > ELEVENLABS_API_KEY=your-api-key > ``` > > If `.env` already has an `ELEVENLABS_API_KEY=...` line, replace that line. > Tell me when you've saved it, without sharing the key. Then wait for the user to confirm that the key is saved locally. ### Step 2: Validate and configure After the user says the key is saved: 1. Re-check both `.env` and the current environment for `ELEVENLABS_API_KEY`, but treat `.env` as the source of truth for this step. 2. If `.env` contains a value, validate that value even when the current environment also has a different `ELEVENLABS_API_KEY`. 3. If `.env` does not contain the key: - Tell the user `.env` does not appear to contain `ELEVENLABS_API_KEY`. - Show the expected line again. - If the current environment does contain a key, note that this step still requires saving the key in `.env`. - Remind them not to paste the key into chat. 4. If a `.env` key is found, validate it: ```text GET https://api.elevenlabs.io/v1/user Header: xi-api-key: <local-api-key> ``` 5. If validation fails: - Tell the user the local key appears invalid or expired. - Remind them of the API keys page. - Ask them to replace the `.env` value and tell you when it is saved. 6. If validation succeeds, confirm: > Done. ElevenLabs is configured and the key in `.env` works. ## Safety Rules - Never ask the user to paste an API key, token, or secret into chat. - Never print or echo API key values from environment variables or `.env`. - Prefer `.env` or managed secrets over shell history for persistent local configuration. - For browser or client-side apps, keep `ELEVENLABS_API_KEY` on the server and issue short-lived tokens where applicable.
sound-effects3.77 KB
---
name: sound-effects
description: Generate sound effects from text descriptions using ElevenLabs. Use when creating sound effects, generating audio textures, producing ambient sounds, cinematic impacts, UI sounds, or any audio that isn't speech. Supports looping, duration control, and prompt influence tuning.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Sound Effects
Generate sound effects from text descriptions — supports looping, custom duration, and prompt adherence control.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio = client.text_to_sound_effects.convert(
text="Thunder rumbling in the distance with light rain",
)
with open("thunder.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
const client = new ElevenLabsClient();
const audio = await client.textToSoundEffects.convert({
text: "Thunder rumbling in the distance with light rain",
});
audio.pipe(createWriteStream("thunder.mp3"));
```
### CLI
```bash
elevenlabs text-to-sound-effects convert \
--text "Thunder rumbling in the distance with light rain" \
--output thunder.mp3
```
## Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `text` | string (required) | — | Description of the desired sound effect |
| `model_id` | string | `eleven_text_to_sound_v2` | Model to use |
| `duration_seconds` | number \| null | null (auto) | Duration 0.5–30s; auto-calculated if null |
| `prompt_influence` | number \| null | 0.3 | How closely to follow the prompt (0–1) |
| `loop` | boolean | false | Generate a seamlessly looping sound (v2 model only) |
## Examples with Parameters
```python
# Looping ambient sound, 10 seconds
audio = client.text_to_sound_effects.convert(
text="Gentle forest ambiance with birds chirping",
duration_seconds=10.0,
prompt_influence=0.5,
loop=True,
)
# Short UI sound, high prompt adherence
audio = client.text_to_sound_effects.convert(
text="Soft notification chime",
duration_seconds=1.0,
prompt_influence=0.8,
)
```
## Output Formats
Pass `--output-format` (CLI) or `output_format` as an SDK parameter:
| Format | Description |
|--------|-------------|
| `mp3_44100_128` | MP3 44.1kHz 128kbps (default) |
| `pcm_44100` | Raw uncompressed CD quality |
| `opus_48000_128` | Opus 48kHz 128kbps — efficient compressed |
| `ulaw_8000` | μ-law 8kHz — telephony |
Full list: `mp3_22050_32`, `mp3_24000_48`, `mp3_44100_32`, `mp3_44100_64`, `mp3_44100_96`, `mp3_44100_128`, `mp3_44100_192`, `pcm_8000`, `pcm_16000`, `pcm_22050`, `pcm_24000`, `pcm_32000`, `pcm_44100`, `pcm_48000`, `ulaw_8000`, `alaw_8000`, `opus_48000_32`, `opus_48000_64`, `opus_48000_96`, `opus_48000_128`, `opus_48000_192`.
## Prompt Tips
- Be specific: "Heavy rain on a tin roof" > "Rain"
- Combine elements: "Footsteps on gravel with distant traffic"
- Specify style: "Cinematic braam, horror" or "8-bit retro jump sound"
- Mention mood/context: "Eerie wind howling through an abandoned building"
## Error Handling
```python
try:
audio = client.text_to_sound_effects.convert(text="Explosion")
except Exception as e:
print(f"API error: {e}")
```
Common errors:
- **401**: Invalid API key
- **422**: Invalid parameters (check duration range, prompt_influence range)
- **429**: Rate limit exceeded
## References
- [Installation Guide](references/installation.md)
Referenced files: 1
speech-engine10.1 KB
---
name: speech-engine
description: Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Use when building Speech Engine servers, WebSocket handlers, WebRTC browser clients, conversation token endpoints, interruption-aware streaming responses, or voice-enabled chat agents that connect developer-owned server logic to ElevenLabs speech-to-text and text-to-speech.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Speech Engine
Add a real-time voice interface to a custom agent. ElevenLabs handles microphone audio, speech-to-text, turn-taking, text-to-speech, and browser playback; your server exposes a Speech Engine WebSocket endpoint and streams response text back.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only. For deeper SDK details, read [JavaScript SDK Reference](references/javascript-sdk-reference.md) or [Python SDK Reference](references/python-sdk-reference.md).
## When to Use
Use Speech Engine when the user wants to:
- Add voice to an existing chat app or custom server pipeline
- Add voice to OpenClaw, Hermes, or a similar agent runtime while keeping agent logic on the developer-owned server
- Build a developer-hosted WebSocket server for ElevenLabs voice conversations
- Stream response text back as spoken audio after your server validates user intent
- Handle user interruptions while a response is still streaming
- Build a browser client with `@elevenlabs/react` or `@elevenlabs/client` using a server-issued conversation token
Use the `agents` skill instead when the user is creating or configuring a hosted ElevenLabs Conversational AI agent with platform-managed prompts, tools, workflows, phone numbers, or widgets.
## How It Works
Each Speech Engine WebSocket connection represents one conversation.
1. The browser sends user audio to ElevenLabs.
2. ElevenLabs sends speech-recognition events to your server.
3. Your server derives trusted application state without letting raw speech text control tools or privileged actions.
4. Your server streams text back through the SDK.
5. ElevenLabs converts the response to speech and plays it in the browser.
The SDK manages WebSocket routing, request verification, session lifecycle, ping/pong, turn-taking, and interruption handling. `sendResponse()` / `send_response()` accepts a string or async iterable of response text.
Treat speech-recognition text as untrusted user input. Do not map raw speech text directly into model roles, responses, or tool calls. Use deterministic validation, allowlisted intents, or explicit user confirmation before any transcript-derived value affects downstream response or tool logic.
## Implementation Flow
1. Install server dependencies and configure `ELEVENLABS_API_KEY`.
2. Expose your Speech Engine server through a public HTTPS URL for local development, for example with `ngrok http 3001`.
3. Create a Speech Engine resource with `ws_url` / `wsUrl` pointing at the public WebSocket URL, usually `wss://.../ws`.
4. Store the returned Speech Engine ID, for example in `ELEVENLABS_SPEECH_ENGINE_ID`.
5. Start a Speech Engine server with `engine.serve(...)` in Python or `speechEngine.attach(...)` in TypeScript.
6. Issue browser conversation tokens from a server endpoint. Never put `ELEVENLABS_API_KEY` in browser code.
7. Start the client session with `conversationToken`; if the agent should greet first, enable the first-message override on the Speech Engine resource, then set `overrides.agent.firstMessage` in the client.
## Create a Speech Engine
### Python
```python
import asyncio
import os
from dotenv import load_dotenv
from elevenlabs import AsyncElevenLabs
load_dotenv()
elevenlabs = AsyncElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
async def main():
engine = await elevenlabs.speech_engine.create(
name="My Speech Engine",
speech_engine={"ws_url": os.environ["PUBLIC_WS_URL"]},
overrides={"first_message": True},
)
print(engine.engine_id)
asyncio.run(main())
```
### TypeScript
```typescript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const elevenlabs = new ElevenLabsClient({
apiKey: process.env.ELEVENLABS_API_KEY,
});
const engine = await elevenlabs.speechEngine.create({
name: "My Speech Engine",
speechEngine: { wsUrl: process.env.PUBLIC_WS_URL! },
overrides: { firstMessage: true },
});
console.log(engine.engineId);
```
`PUBLIC_WS_URL` should look like `wss://example.ngrok.app/ws` locally or your production WebSocket route in deployment.
The create request can also configure `tts`, `asr`, `turn`, `speech_engine.request_headers` / `speechEngine.requestHeaders`, `overrides`, and `privacy` for custom voices, transcription keywords, turn-taking, server auth headers, client-provided first messages, and recording behavior. See the SDK reference files for expanded examples.
## Server Pattern
Run the Speech Engine server at the `ws_url` / `wsUrl` configured on the resource. Keep response generation behind your own validation boundary: raw speech-recognition text should not directly control responses, tools, secrets, or other privileged actions.
### Python
```python
engine = await elevenlabs.speech_engine.get(os.environ["ELEVENLABS_SPEECH_ENGINE_ID"])
await engine.serve(port=3001, path="/ws", debug=True, callbacks=validated_callbacks)
```
### TypeScript
```typescript
const engine = await elevenlabs.speechEngine.get(process.env.ELEVENLABS_SPEECH_ENGINE_ID!);
engine.attach(httpServer, "/ws", { debug: true, ...validatedCallbacks });
```
In TypeScript, pass interruption signals to downstream async work when it supports cancellation so interrupted responses stop quickly. In Python, the SDK cancels the previous turn handler when a newer turn arrives.
Server callbacks can distinguish clean closes from dropped connections: use `onClose` / `on_close` for clean disconnects and `onDisconnect` / `on_disconnect` for unexpected WebSocket drops.
Security note: speech-recognition text can contain prompt-injection attempts from user speech or played audio. Treat it as untrusted input. Convert it into trusted application state before invoking response generation, tools, or privileged workflows.
### Disabling authentication (advanced, dangerous)
Both `engine.attach()` (TypeScript) and `engine.serve()` / `SpeechEngineServer` (Python) verify a JWT on every incoming WebSocket by default. This is what proves the connection is really coming from ElevenLabs and not from an attacker who guessed the URL. **Do not turn this off.**
An escape hatch exists — `disableAuth: true` in the callback options (TypeScript) or `disable_auth=True` on `serve()` / `SpeechEngineServer(...)` (Python) — for the narrow case where a compensating network-level control is already in place. Without such a control, disabling auth means **any client on the internet that finds your URL can open sessions**. Concretely, an attacker can:
- open unlimited conversations to drain your ElevenLabs quota and downstream LLM budget
- feed crafted transcripts to your response pipeline, effectively impersonating a user
- use your server as an oracle to probe backend state, tools, or prompts
Only recommend `disableAuth` / `disable_auth` when the user has already implemented **at least one** of:
- **IP allowlist** — the server (or an upstream firewall / load balancer / API gateway) only accepts inbound traffic from [ElevenLabs' documented egress ranges](https://elevenlabs.io/docs/overview/capabilities/speech-engine#ip-allowlisting).
- **Custom shared-secret header** — a secret header configured on the Speech Engine resource via `speech_engine.request_headers` / `speechEngine.requestHeaders` at create time, validated by an upstream proxy (or by the developer's own middleware in front of `attach()` / `serve()`) before requests reach the SDK.
If the user cannot confirm one of the above is in place, leave the default authentication on. Skipping JWT verification without a mitigation is not an optimization or a convenience — it is unauthenticated public compute.
## Browser Client
Create a server-side token endpoint and have the browser request a token before starting the microphone session. Keep the Speech Engine ID and API key on the server. If the client passes `overrides.agent.firstMessage`, the Speech Engine resource must have the first-message override enabled.
```typescript
import express from "express";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const app = express();
const elevenlabs = new ElevenLabsClient();
app.get("/api/token", async (_req, res) => {
const response = await elevenlabs.conversationalAi.conversations.getWebrtcToken({
agentId: process.env.ELEVENLABS_SPEECH_ENGINE_ID!,
});
res.json({ token: response.token });
});
```
React clients can use `@elevenlabs/react`:
```tsx
import { useConversation } from "@elevenlabs/react";
export function VoiceControls() {
const conversation = useConversation({
onConnect: () => console.log("connected"),
onDisconnect: () => console.log("disconnected"),
onError: (error) => console.error(error),
});
async function startConversation() {
await navigator.mediaDevices.getUserMedia({ audio: true });
const { token } = await fetch("/api/token").then((res) => res.json());
await conversation.startSession({
conversationToken: token,
overrides: {
agent: { firstMessage: "Hello! How can I help you today?" },
},
});
}
return <button onClick={startConversation}>Start conversation</button>;
}
```
If a WebRTC browser session stalls or logs `/rtc/v1` 404s, `v1 RTC path not found`, or `could not establish pc connection`, pin `livekit-client` to `2.16.1` in the app's `package.json` until the upstream LiveKit compatibility issue is resolved:
```json
{
"overrides": {
"livekit-client": "2.16.1"
}
}
```
## References
- [Installation Guide](references/installation.md)
- [JavaScript SDK Reference](references/javascript-sdk-reference.md)
- [Python SDK Reference](references/python-sdk-reference.md)
Referenced files: 3
speech-to-text9.67 KB
---
name: speech-to-text
description: Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Speech-to-Text
Transcribe audio to text with Scribe v2 - supports 90+ languages, speaker diarization, and word-level timestamps.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
with open("audio.mp3", "rb") as audio_file:
result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
print(result.text)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream } from "fs";
const client = new ElevenLabsClient();
const result = await client.speechToText.convert({
file: createReadStream("audio.mp3"),
modelId: "scribe_v2",
});
console.log(result.text);
```
### CLI
```bash
elevenlabs speech-to-text convert --file audio.mp3 --model-id scribe_v2
```
## Models
| Model ID | Description | Best For |
|----------|-------------|----------|
| `scribe_v2` | State-of-the-art accuracy, 90+ languages | Batch transcription, subtitles, long-form audio |
| `scribe_v2_realtime` | Low latency (~150ms) | Live transcription, voice agents |
| `scribe_v2_realtime_turbo` | Realtime transcription variant | Live transcription |
| `scribe_v2_realtime_lite` | Realtime transcription variant | Live transcription |
## Transcription with Timestamps
Word-level timestamps include type classification and speaker identification:
```python
result = client.speech_to_text.convert(
file=audio_file, model_id="scribe_v2", timestamps_granularity="word"
)
for word in result.words:
print(f"{word.text}: {word.start}s - {word.end}s (type: {word.type})")
```
## Speaker Diarization
Identify WHO said WHAT - the model labels each word with a speaker ID, useful for meetings, interviews, or any multi-speaker audio:
```python
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
diarize=True
)
for word in result.words:
print(f"[{word.speaker_id}] {word.text}")
```
For call recordings, the batch API can label diarized speakers as `agent` and `customer` by setting `detect_speaker_roles=true` alongside `diarize=true`. This option is not compatible with `use_multi_channel=true`.
If your workspace has registered speaker profiles, set `use_speaker_library=true` with `diarize=true` to match detected speakers against the speaker library.
```bash
elevenlabs speech-to-text convert \
--file call.mp3 \
--model-id scribe_v2 \
--diarize true \
--detect-speaker-roles true \
--use-speaker-library true
```
## Multichannel Audio
Use `use_multi_channel=true` when each speaker is isolated on a separate audio channel. By default, the API returns one transcript per channel under `transcripts`; set `multichannel_output_style="combined"` to receive one transcript merged by timestamp, with `channel_index` on each word.
```python
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
use_multi_channel=True,
multichannel_output_style="combined",
)
```
## Keyterm Prompting
Help the model recognize specific words it might otherwise mishear - product names, technical jargon, or unusual spellings (up to 100 terms):
```python
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
keyterms=["ElevenLabs", "Scribe", "API"]
)
```
## Language Detection
Automatic detection with optional language hint:
```python
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
language_code="eng" # ISO 639-1 or ISO 639-3 code
)
print(f"Detected: {result.language_code} ({result.language_probability:.0%})")
```
## Supported Formats
**Audio:** MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus
**Video:** MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP
**Limits:** Up to 5.0GB file size, 10 hours duration
## Response Format
```json
{
"text": "The full transcription text",
"language_code": "eng",
"language_probability": 0.98,
"words": [
{"text": "The", "start": 0.0, "end": 0.15, "type": "word", "speaker_id": "speaker_0"},
{"text": " ", "start": 0.15, "end": 0.16, "type": "spacing", "speaker_id": "speaker_0"}
]
}
```
**Word types:**
- `word` - An actual spoken word
- `spacing` - Whitespace between words (useful for precise timing)
- `audio_event` - Non-speech sounds the model detected (laughter, applause, music, etc.)
## Error Handling
```python
try:
result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
except Exception as e:
print(f"Transcription failed: {e}")
```
Common errors:
- **401**: Invalid API key
- **422**: Invalid parameters
- **429**: Rate limit exceeded
## Tracking Costs
Monitor usage via `request-id` response header:
```python
response = client.speech_to_text.with_raw_response.convert(file=audio_file, model_id="scribe_v2")
result = response.data
print(f"Request ID: {response.headers.get('request-id')}")
```
## Real-Time Streaming
For live transcription with ultra-low latency (~150ms), use the real-time API. The real-time API produces two types of transcripts:
- **Partial transcripts**: Interim results that update frequently as audio is processed - use these for live feedback (e.g., showing text as the user speaks)
- **Committed transcripts**: Final, stable results after you "commit" - use these as the source of truth for your application
A "commit" tells the model to finalize the current segment. You can commit manually (e.g., when the user pauses) or use Voice Activity Detection (VAD) to auto-commit on silence.
### Python (Server-Side)
```python
import asyncio
from elevenlabs import ElevenLabs
client = ElevenLabs()
async def transcribe_realtime():
async with client.speech_to_text.realtime.connect(
model_id="scribe_v2_realtime",
include_timestamps=True,
keyterms=["ElevenLabs", "Scribe"],
no_verbatim=True,
) as connection:
await connection.stream_url("https://example.com/audio.mp3")
async for event in connection:
if event.type == "partial_transcript":
print(f"Partial: {event.text}")
elif event.type == "committed_transcript":
print(f"Final: {event.text}")
asyncio.run(transcribe_realtime())
```
### JavaScript (Client-Side with React)
```typescript
import { useScribe, CommitStrategy } from "@elevenlabs/react";
function TranscriptionComponent() {
const [transcript, setTranscript] = useState("");
const scribe = useScribe({
modelId: "scribe_v2_realtime",
commitStrategy: CommitStrategy.VAD, // Auto-commit on silence for mic input
keyterms: ["ElevenLabs", "Scribe"],
noVerbatim: true,
includeLanguageDetection: true,
onPartialTranscript: (data) => console.log("Partial:", data.text),
onCommittedTranscript: (data) => setTranscript((prev) => prev + data.text),
});
const start = async () => {
// Get token from your backend (never expose API key to client)
const { token } = await fetch("/scribe-token").then((r) => r.json());
await scribe.connect({
token,
microphone: { echoCancellation: true, noiseSuppression: true },
});
};
return <button onClick={start}>Start Recording</button>;
}
```
### Commit Strategies
| Strategy | Description |
|----------|-------------|
| **Manual** | You call `commit()` when ready - use for file processing or when you control the audio segments |
| **VAD** | Voice Activity Detection auto-commits when silence is detected - use for live microphone input |
Set `includeLanguageDetection: true` to receive the detected language code in delayed final
transcript events.
```typescript
// React: set commitStrategy on the hook (recommended for mic input)
import { useScribe, CommitStrategy } from "@elevenlabs/react";
const scribe = useScribe({
modelId: "scribe_v2_realtime",
commitStrategy: CommitStrategy.VAD,
keyterms: ["ElevenLabs", "Scribe"],
noVerbatim: true,
// Optional VAD tuning:
vadSilenceThresholdSecs: 1.5,
vadThreshold: 0.4,
});
```
```javascript
// JavaScript client: pass vad config on connect
const connection = await client.speechToText.realtime.connect({
modelId: "scribe_v2_realtime",
keyterms: ["ElevenLabs", "Scribe"],
noVerbatim: true,
vad: {
silenceThresholdSecs: 1.5,
threshold: 0.4,
},
});
```
### Event Types
| Event | Description |
|-------|-------------|
| `partial_transcript` | Live interim results |
| `final_transcript` | Stable segment result sent before the segment is committed |
| `final_transcript_with_timestamps` | Delayed final result with timestamps and/or detected language |
| `committed_transcript` | Final results after commit |
| `committed_transcript_with_timestamps` | Final with word timing |
| `committed_transcript_entities` | Entities detected in a committed segment |
| `invalid_request` | Connection parameters were rejected and the session closes |
| `error` | Error occurred |
See real-time references for complete documentation.
## References
- [Installation Guide](references/installation.md)
- [Transcription Options](references/transcription-options.md)
- [Real-Time Client-Side Streaming](references/realtime-client-side.md)
- [Real-Time Server-Side Streaming](references/realtime-server-side.md)
- [Commit Strategies](references/realtime-commit-strategies.md)
- [Real-Time Event Reference](references/realtime-events.md)
Referenced files: 6
text-to-speech7.32 KB
---
name: text-to-speech
description: Convert text to speech using ElevenLabs voice AI. Use when generating audio from text, creating voiceovers, building voice apps, or synthesizing speech in 70+ languages.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Text-to-Speech
Generate natural speech from text - supports 70+ languages, multiple models for quality vs latency tradeoffs.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio = client.text_to_speech.convert(
text="Hello, welcome to ElevenLabs!",
voice_id="JBFqnCBsd6RMkjVDRZzb", # George
model_id="eleven_multilingual_v2"
)
with open("output.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createWriteStream } from "fs";
import { Readable } from "stream";
const client = new ElevenLabsClient();
const audio = await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
text: "Hello, welcome to ElevenLabs!",
modelId: "eleven_multilingual_v2",
});
// convert() returns a web ReadableStream — bridge it to a Node stream to write to disk
Readable.fromWeb(audio).pipe(createWriteStream("output.mp3"));
```
### CLI
```bash
elevenlabs text-to-speech convert --voice-id JBFqnCBsd6RMkjVDRZzb \
--text "Hello!" --model-id eleven_multilingual_v2 --output output.mp3
```
The CLI reads `ELEVENLABS_API_KEY` from the environment automatically.
## Models
| Model ID | Languages | Latency | Best For |
|----------|-----------|---------|----------|
| `eleven_v3` | 70+ | Standard | Highest quality, emotional range |
| `eleven_multilingual_v2` | 29 | Standard | High quality, long-form content |
| `eleven_flash_v2_5` | 32 | ~75ms | Ultra-low latency, real-time |
| `eleven_flash_v2` | English | ~75ms | English-only, fastest |
| `eleven_turbo_v2_5` | 32 | ~250-300ms | Balanced quality/speed |
| `eleven_turbo_v2` | English | ~250-300ms | English-only, balanced |
## Voice IDs
Use pre-made voices or create custom voices in the dashboard.
**Popular voices:**
- `JBFqnCBsd6RMkjVDRZzb` - George (male, narrative)
- `EXAVITQu4vr4xnSDxMaL` - Sarah (female, soft)
- `onwK4e9ZLuTAKqWW03F9` - Daniel (male, authoritative)
- `XB0fDUnXU5powFXDhCwa` - Charlotte (female, conversational)
```python
voices = client.voices.get_all()
for voice in voices.voices:
print(f"{voice.voice_id}: {voice.name}")
```
## Voice Settings
Fine-tune how the voice sounds:
- **Stability**: How consistent the voice stays. Lower values = more emotional range and variation, but can sound unstable. Higher = steady, predictable delivery.
- **Similarity boost**: How closely to match the original voice sample. Higher values sound more like the original but may amplify audio artifacts.
- **Style**: Exaggerates the voice's unique style characteristics (only works with v2+ models).
- **Speaker boost**: Post-processing that enhances clarity and voice similarity.
```python
from elevenlabs import VoiceSettings
audio = client.text_to_speech.convert(
text="Customize my voice settings.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
voice_settings=VoiceSettings(
stability=0.5,
similarity_boost=0.75,
style=0.5,
speed=1.0, # 0.25 to 4.0 (default 1.0)
use_speaker_boost=True
)
)
```
## Language Selection
Use `language_code` with models that support language enforcement to guide pronunciation and text normalization. Unsupported language codes are ignored, and `language_code` is not supported on `eleven_multilingual_v2`.
```python
audio = client.text_to_speech.convert(
text="Bonjour, comment allez-vous?",
voice_id="JBFqnCBsd6RMkjVDRZzb",
model_id="eleven_v3",
language_code="fr" # ISO 639-1 code
)
```
## Text Normalization
Controls how numbers, dates, and abbreviations are converted to spoken words. For example, "01/15/2026" becomes "January fifteenth, twenty twenty-six":
- `"auto"` (default): Model decides based on context
- `"on"`: Always normalize (use when you want natural speech)
- `"off"`: Speak literally (use when you want "zero one slash one five...")
```python
audio = client.text_to_speech.convert(
text="Call 1-800-555-0123 on 01/15/2026",
voice_id="JBFqnCBsd6RMkjVDRZzb",
apply_text_normalization="on"
)
```
## Request Stitching
When generating long audio in multiple requests, the audio can have pops, unnatural pauses, or tone shifts at the boundaries. Request stitching solves this by letting each request know what comes before/after it:
```python
# First request
audio1 = client.text_to_speech.convert(
text="This is the first part.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
next_text="And this continues the story."
)
# Second request using previous context
audio2 = client.text_to_speech.convert(
text="And this continues the story.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
previous_text="This is the first part."
)
```
## Output Formats
| Format | Description |
|--------|-------------|
| `mp3_44100_128` | MP3 44.1kHz 128kbps (default) - compressed, good for web/apps |
| `mp3_44100_192` | MP3 44.1kHz 192kbps (Creator+) - higher quality compressed |
| `mp3_44100_64` | MP3 44.1kHz 64kbps - lower quality, smaller files |
| `mp3_22050_32` | MP3 22.05kHz 32kbps - smallest MP3 files |
| `pcm_16000` | Raw PCM 16kHz - use for real-time processing |
| `pcm_22050` | Raw PCM 22.05kHz |
| `pcm_24000` | Raw PCM 24kHz - good balance for streaming |
| `pcm_44100` | Raw PCM 44.1kHz (Pro+) - CD quality |
| `pcm_48000` | Raw PCM 48kHz (Pro+) - highest quality |
| `ulaw_8000` | μ-law 8kHz - standard for phone systems (Twilio, telephony) |
| `alaw_8000` | A-law 8kHz - telephony (alternative to μ-law) |
| `opus_48000_64` | Opus 48kHz 64kbps - efficient streaming codec |
| `wav_44100` | WAV 44.1kHz - uncompressed with headers |
## Streaming
For real-time applications, use the `stream` method (returns audio chunks as they're generated):
```python
audio_stream = client.text_to_speech.stream(
text="This text will be streamed as audio.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
model_id="eleven_flash_v2_5" # Ultra-low latency
)
for chunk in audio_stream:
play_audio(chunk)
```
See [references/streaming.md](references/streaming.md) for WebSocket streaming.
## Error Handling
```python
try:
audio = client.text_to_speech.convert(
text="Generate speech",
voice_id="invalid-voice-id"
)
except Exception as e:
print(f"API error: {e}")
```
Common errors:
- **401**: Invalid API key
- **422**: Invalid parameters (check voice_id, model_id)
- **429**: Rate limit exceeded
## Tracking Costs
Monitor character usage via response headers (`x-character-count`, `request-id`):
```python
response = client.text_to_speech.convert.with_raw_response(
text="Hello!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2"
)
audio = response.parse()
print(f"Characters used: {response.headers.get('x-character-count')}")
```
## References
- [Installation Guide](references/installation.md)
- [Streaming Audio](references/streaming.md)
- [Voice Settings](references/voice-settings.md)
Referenced files: 3
voice-changer11.6 KB
---
name: voice-changer
description: Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API. Use when converting one voice to another, changing the speaker/narrator of an existing recording, dubbing a voice-over in a different voice, creating character voices from a scratch performance, anonymizing a speaker, or any "voice conversion / voice transfer / speech-to-speech" task. Make sure to use this skill whenever the user mentions voice changing, voice conversion, speech-to-speech, swapping a voice in audio, re-voicing a clip, or applying a different voice to an existing recording — even if they don't explicitly say "voice changer". Do not use this skill to create or clone a new voice from a voice sample — that is voice cloning (IVC/PVC), a separate feature; this skill only converts an existing recording into an existing voice_id.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Voice Changer
Transform the voice in an audio recording into a different target voice. Voice Changer (previously called Speech-to-Speech — the API endpoint and SDK methods still use the `speech_to_speech` / `speechToSpeech` name) keeps the original performance — emotion, pacing, intonation, breaths, whispers, laughs, cries — and only swaps who is speaking.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
## Key Facts
- **Maximum input length:** 5 minutes per request — split longer recordings into chunks and stitch the outputs.
- **Maximum file size:** 50 MB per request — compress to MP3 if your source is larger.
- **Pricing:** 1,000 characters per minute of audio processed (duration-based, not text-based).
- **Recommended model:** `eleven_multilingual_sts_v2` — often outperforms `eleven_english_sts_v2` even for English-only content.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
with open("source.mp3", "rb") as audio_file:
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb", # George
audio=audio_file,
model_id="eleven_multilingual_sts_v2",
output_format="mp3_44100_128",
)
with open("converted.mp3", "wb") as f:
for chunk in audio_stream:
f.write(chunk)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";
const client = new ElevenLabsClient();
const audioStream = await client.speechToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
audio: createReadStream("source.mp3"),
modelId: "eleven_multilingual_sts_v2",
outputFormat: "mp3_44100_128",
});
audioStream.pipe(createWriteStream("converted.mp3"));
```
### CLI
```bash
elevenlabs speech-to-speech convert \
--voice-id JBFqnCBsd6RMkjVDRZzb \
--audio source.mp3 \
--model-id eleven_multilingual_sts_v2 \
--output-format mp3_44100_128 \
--output converted.mp3
```
## Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `voice_id` | string (required) | — | Target voice to speak in. Use a pre-made voice ID, a cloned voice, or a voice from the library |
| `audio` | file (required) | — | Source audio whose performance (emotion, timing, delivery) will be preserved |
| `model_id` | string | `eleven_english_sts_v2` | `eleven_multilingual_sts_v2` for 29 languages, `eleven_english_sts_v2` for English-only |
| `output_format` | string | `mp3_44100_128` | See output formats table below |
| `voice_settings` | JSON string | — | Override stored voice settings for this request only |
| `seed` | integer | — | Best-effort deterministic sampling (0 – 4294967295) |
| `remove_background_noise` | boolean | `false` | Run the isolation model on the input before conversion |
| `file_format` | string | `other` | `other` for any encoded audio, or `pcm_s16le_16` for 16-bit PCM mono @ 16kHz little-endian (lower latency) |
| `optimize_streaming_latency` | int (query) | — | 0–4. Trade quality for latency. `4` is fastest but disables the text normalizer |
| `enable_logging` | boolean (query) | `true` | Set to `false` for zero-retention mode (enterprise only — disables history/stitching) |
## Models
| Model ID | Languages | Best For |
|----------|-----------|----------|
| `eleven_multilingual_sts_v2` | 29 | Recommended for everything — often outperforms the English model even on English audio |
| `eleven_english_sts_v2` | English | API default — English-only fallback |
Only models whose `can_do_voice_conversion` property is true can be used here. Voice Changer does not currently have a low-latency "flash/turbo" tier — if you need one, keep `pcm_s16le_16` input, an `opus_*` / low-bitrate `mp3_*` output, and raise `optimize_streaming_latency`.
### Languages (`eleven_multilingual_sts_v2`)
English (US, UK, AU, CA), Japanese, Chinese, German, Hindi, French (FR, CA), Korean, Portuguese (BR, PT), Italian, Spanish (ES, MX), Indonesian, Dutch, Turkish, Filipino, Polish, Swedish, Bulgarian, Romanian, Arabic (SA, AE), Czech, Greek, Finnish, Croatian, Malay, Slovak, Danish, Tamil, Ukrainian, Russian.
## Target Voices
Use any voice ID from pre-made voices, your cloned voices, or the voice library.
**Popular voices:**
- `JBFqnCBsd6RMkjVDRZzb` — George (male, narrative)
- `EXAVITQu4vr4xnSDxMaL` — Sarah (female, soft)
- `onwK4e9ZLuTAKqWW03F9` — Daniel (male, authoritative)
- `XB0fDUnXU5powFXDhCwa` — Charlotte (female, conversational)
```python
voices = client.voices.get_all()
for voice in voices.voices:
print(f"{voice.voice_id}: {voice.name}")
```
## Converting from a URL
```python
import requests
from io import BytesIO
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio_url = "https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
response = requests.get(audio_url)
audio_data = BytesIO(response.content)
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=audio_data,
model_id="eleven_multilingual_sts_v2",
output_format="mp3_44100_128",
)
with open("converted.mp3", "wb") as f:
for chunk in audio_stream:
f.write(chunk)
```
## Voice Settings Override
Fine-tune the target voice for a single request without changing its stored defaults:
```python
from elevenlabs import VoiceSettings
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=audio_file,
model_id="eleven_multilingual_sts_v2",
voice_settings=VoiceSettings(
stability=0.5,
similarity_boost=0.75,
style=0.0,
use_speaker_boost=True,
),
)
```
- **Stability**: lower = more emotional range (follows the source more freely), higher = steadier delivery.
- **Similarity boost**: higher = closer to the target voice's timbre, may amplify source artifacts.
- **Style**: exaggerates the target voice's unique characteristics (v2+ models).
- **Speaker boost**: post-processing to sharpen clarity of the target voice.
## Cleaning Up Noisy Source Audio
If the input recording is noisy, either pre-process with the voice-isolator skill or pass `remove_background_noise=True` to do it in a single call:
```python
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=audio_file,
model_id="eleven_multilingual_sts_v2",
remove_background_noise=True,
)
```
Cleaner input almost always produces better conversion — the model is trying to match phonemes and prosody, and background noise gets in the way.
## Low-Latency PCM Input
If you already have raw 16-bit PCM mono @ 16kHz, passing `file_format="pcm_s16le_16"` skips decoding and reduces latency:
```python
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=pcm_bytes,
model_id="eleven_multilingual_sts_v2",
file_format="pcm_s16le_16",
)
```
Pair this with `optimize_streaming_latency` (0–4) as a query param for further latency reductions at some quality cost.
## Output Formats
| Format | Description |
|--------|-------------|
| `mp3_44100_128` | MP3 44.1kHz 128kbps (default) — good for web/apps |
| `mp3_44100_192` | MP3 44.1kHz 192kbps (Creator+) — higher quality |
| `mp3_44100_64` | MP3 44.1kHz 64kbps — smaller files |
| `mp3_22050_32` | MP3 22.05kHz 32kbps — smallest MP3 |
| `pcm_16000` | Raw PCM 16kHz — real-time pipelines |
| `pcm_24000` | Raw PCM 24kHz — good streaming balance |
| `pcm_44100` | Raw PCM 44.1kHz (Pro+) — CD quality |
| `pcm_48000` | Raw PCM 48kHz (Pro+) — highest quality |
| `ulaw_8000` | μ-law 8kHz — Twilio / telephony |
| `alaw_8000` | A-law 8kHz — telephony |
| `opus_48000_64` | Opus 48kHz 64kbps — efficient streaming |
## Deterministic Output
Pass a `seed` to make repeated conversions of the same input return (best-effort) identical audio — useful for testing and A/B comparisons.
```python
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=audio_file,
model_id="eleven_multilingual_sts_v2",
seed=12345,
)
```
## Input Audio Best Practices
The conversion quality is bounded by the input recording — the model can only swap the timbre, not rescue a bad source. A few practical rules:
- **Be expressive.** Whisper, shout, laugh, cry — the model preserves all of it. Flat input gives you flat output.
- **Watch microphone gain.** Too quiet and the model under-detects phonemes; too loud and clipping bleeds into the conversion. Aim for healthy peaks, no clipping.
- **Accent and cadence transfer from the source, not the target.** If you read in an American accent and target the British "George" voice, you get George's timbre with an American accent. To dub *into* a different accent or language, record someone speaking in that target accent/language and convert into a cloned/library voice.
- **Clean up noise first.** Either pass `remove_background_noise=True` or run the source through the voice-isolator skill before conversion. Noise hurts more here than in TTS.
- **Split long recordings.** Anything over 5 minutes must be chunked. Cut at natural pauses, convert each piece, and concatenate the resulting audio.
## Common Workflows
- **Re-voice a narration** — keep the performance of a scratch recording, swap in a different narrator voice.
- **Localize / dub** — convert a voice-over into the same speaker's cloned voice in another language (using `eleven_multilingual_sts_v2`).
- **Create character voices** — act out a line yourself, convert into a distinctive character voice for games or animation.
- **Anonymize a speaker** — replace a recognizable voice with a neutral pre-made voice while preserving what was said and how.
- **Pair with voice-isolator** — isolate the source voice first (or set `remove_background_noise=True`) for noisy field recordings before conversion.
- **Pair with voice cloning** — clone a target voice from a short sample, then use its `voice_id` here as the conversion target.
## Error Handling
```python
try:
audio_stream = client.speech_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb",
audio=audio_file,
model_id="eleven_multilingual_sts_v2",
)
except Exception as e:
print(f"Voice changer failed: {e}")
```
Common errors:
- **401**: Invalid API key
- **422**: Invalid parameters (check `voice_id`, `model_id`, or `file_format` vs the supplied audio)
- **429**: Rate limit exceeded
## References
- [Installation Guide](references/installation.md)
Referenced files: 1
voice-isolator3.61 KB
---
name: voice-isolator
description: Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API. Use when cleaning up noisy recordings, removing music or background ambience from dialogue, isolating speech from field recordings, preparing audio for transcription, extracting vocals, or any "denoise / clean up / isolate voice" task.
license: MIT
compatibility: Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
metadata: {"openclaw": {"requires": {"env": ["ELEVENLABS_API_KEY"]}, "primaryEnv": "ELEVENLABS_API_KEY"}}
---
# ElevenLabs Voice Isolator
Removes background noise from audio and isolates vocals/speech — useful for cleaning up noisy recordings, prepping audio for transcription, or pulling dialogue out of a mixed track.
> **Setup:** See [Installation Guide](references/installation.md). For JavaScript, use `@elevenlabs/*` packages only.
## Quick Start
### Python
```python
from elevenlabs import ElevenLabs
client = ElevenLabs()
with open("noisy.mp3", "rb") as audio_file:
audio_stream = client.audio_isolation.convert(audio=audio_file)
with open("clean.mp3", "wb") as f:
for chunk in audio_stream:
f.write(chunk)
```
### JavaScript
```javascript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { createReadStream, createWriteStream } from "fs";
const client = new ElevenLabsClient();
const audioStream = await client.audioIsolation.convert({
audio: createReadStream("noisy.mp3"),
});
audioStream.pipe(createWriteStream("clean.mp3"));
```
### CLI
```bash
elevenlabs audio-isolation convert --audio noisy.mp3 --output clean.mp3
```
## Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `audio` | file (required) | — | Audio file with vocals/speech to isolate |
| `file_format` | string | `other` | `other` for any encoded audio, or `pcm_s16le_16` for 16-bit PCM mono @ 16kHz little-endian (lower latency) |
## Isolating from a URL
```python
import requests
from io import BytesIO
from elevenlabs import ElevenLabs
client = ElevenLabs()
audio_url = "https://example.com/noisy.mp3"
response = requests.get(audio_url)
audio_data = BytesIO(response.content)
audio_stream = client.audio_isolation.convert(audio=audio_data)
with open("clean.mp3", "wb") as f:
for chunk in audio_stream:
f.write(chunk)
```
## Low-Latency PCM Input
If you already have raw 16-bit PCM mono @ 16kHz, passing `file_format="pcm_s16le_16"` skips decoding and reduces latency:
```python
audio_stream = client.audio_isolation.convert(
audio=pcm_bytes,
file_format="pcm_s16le_16",
)
```
## Supported Formats
Any common encoded audio/video container works as input (MP3, WAV, M4A, FLAC, OGG, WebM, MP4, etc.). Response is a streamed MP3 by default.
## Common Workflows
- **Clean up interview/podcast recordings** — strip room tone, HVAC, traffic before editing.
- **Prep noisy audio for Speech-to-Text** — isolate voice first, then pass through `speech_to_text.convert()` for better transcription accuracy.
- **Extract dialogue from mixed tracks** — pull vocals out of a track with music/SFX.
- **Pre-processing for Voice Changer** — isolate the source voice before applying voice transformation.
## Error Handling
```python
try:
audio_stream = client.audio_isolation.convert(audio=audio_file)
except Exception as e:
print(f"Voice isolation failed: {e}")
```
Common errors:
- **401**: Invalid API key
- **422**: Invalid parameters (e.g. wrong `file_format` for the supplied audio)
- **429**: Rate limit exceeded
## References
- [Installation Guide](references/installation.md)
Referenced files: 1
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Eleven Labs Inc.
Package observed Sep 30, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 18:00 UTC
- Collection status
- Collected
plugin_asdk_app_6a8d784b60cc81919aeafbfaeda5fbcf
Download plugin data (JSON)