← ElevenLabsCONTENT HISTORY

Update to ElevenLabs

Snapshot Sep 30, 2026 · 22:53 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "agent-simplification",
  "description": "Simplify a complex ElevenLabs voice agent (typically a large workflow agent) into a leaner architecture — fewer workflow nodes, deduplicated prompts, procedures where they help — while PROVING with a ground-truth test suite that every behaviour of the original is preserved or improved. Use when the user says things like 'simplify this agent', 'flatten this workflow', 'this agent is too complex', 'consolidate these nodes', 'rebuild this agent with fewer nodes', 'can we replace this workflow with procedures', or asks whether a big workflow agent can be made easier to maintain without losing behaviour.",
  "included_files": [],
  "skill_md_contents": "---\nname: agent-simplification\ndescription: \"Simplify a complex ElevenLabs voice agent (typically a large workflow agent) into a leaner architecture — fewer workflow nodes, deduplicated prompts, procedures where they help — while PROVING with a ground-truth test suite that every behaviour of the original is preserved or improved. Use when the user says things like 'simplify this agent', 'flatten this workflow', 'this agent is too complex', 'consolidate these nodes', 'rebuild this agent with fewer nodes', 'can we replace this workflow with procedures', or asks whether a big workflow agent can be made easier to maintain without losing behaviour.\"\n---\n\n# ElevenLabs Agent Simplification\n\nTurn a complex agent (usually a many-node workflow) into a simpler one that is **provably at least\nas good**. The core loop: understand everything the complex agent does → encode it as a test suite\n(ground truth) → build the simpler agent → iterate until it matches or beats the original on every\ntest → A/B on branches → hand over with an honest report.\n\nThe output is never just a simpler agent. It is a simpler agent **plus the test suite that proves\nequivalence**, attached natively in the platform so the client can re-run it.\n\n## Non-negotiable safety rails\n\n- **Never edit the production agent, global tools, the knowledge base, or shared tests.** Work on a\n  duplicate agent or a branch. Tools and KB documents are workspace-global — editing them changes\n  prod. Procedures are per-agent (cloned on duplicate), so a copy's procedures are safe to edit.\n- Before any PATCH, verify you are targeting the copy/branch (`agent_id` / `branch_id` check in the\n  script). Make every write script refuse to run against anything else.\n- Snapshot the full agent JSON before the first change (rollback + later diffing).\n- **PII hygiene:** real transcripts contain names/numbers. Keep them in scratch space, never commit\n  them, and delete them when done. Scrub any dynamic-variable seed derived from a real call\n  (replace the client name, activation date, etc.). Conversation IDs are fine to reference.\n\n## Phase 1 — Understand the complex agent completely\n\nPull the agent JSON (`GET /v1/convai/agents/{id}`) and map every layer:\n\n1. **Base prompt** (`conversation_config.agent.prompt.prompt`).\n2. **Workflow nodes — read `additional_prompt`, not just `prompt.prompt`.** This is the #1 trap:\n   an `override_agent` node's real instructions live in `additional_prompt` (the studio\n   \"Conversation goal\" box, APPENDED to the base prompt). `conversation_config.agent.prompt.prompt`\n   on a node is only set when the \"Override prompt\" toggle is ON. A node whose `prompt.prompt` is\n   empty is NOT an empty node. Extract every node's `additional_prompt` to a file — in a real case\n   16 \"empty-looking\" nodes carried ~100k chars of specialized business logic.\n3. **Edges** — routing conditions (`forward_condition.label`), in/out degrees. The only **dead\n   nodes** are orphans (no incoming edge = unreachable) and fully disconnected nodes (no edges at\n   all). Anything with edges is a live decision point — every node has an LLM choosing among its\n   outgoing edges, even a tool node with an empty tool list — so it cannot be considered for\n   removal unless genuine logic duplication is proven.\n4. **Per-node scoping** — `additional_knowledge_base` / `additional_tool_ids`. Check whether the\n   base agent already exposes the same KB (folders in RAG `auto` mode = every doc) and tools; if so\n   the scoping is redundant, not behaviour.\n5. **Say nodes** — fixed messages the client wants said verbatim (goodbyes, transfer preambles,\n   mandated phrases). These become wording-parity tests later.\n6. **Transfer machinery** — compare workflow `phone_number` node destinations against the base\n   `transfer_to_number` system tool (destination SIP/number + its `condition` text). Often\n   identical → the workflow transfer scaffolding is redundant. Check real usage: count tool fires\n   across recent conversations (which transfer path actually runs in prod?).\n7. **Dispatch tools** (tool nodes in the graph) — these fire **deterministically** whenever the\n   graph reaches that node, whereas agent-attached tools fire at the LLM's discretion. Examine each\n   one: which paths reach it, and what guarantee does it provide? Flattening removes that guarantee,\n   so every dispatch path needs a test proving the tool still fires in the simplified agent. Some\n   dispatch tools also turn out redundant (e.g. a \"say X then transfer\" preamble tool duplicating\n   what the `transfer_to_number` system tool already does) — check real-call tool-fire counts before\n   deciding.\n8. **Procedures, system tools (`built_in_tools`), language config** — note `language_presets` lives\n   at the TOP level of `conversation_config`, not under `.agent`.\n9. **Real conversations** — pull recent calls (`GET /v1/convai/conversations?agent_id=&page_size=100`,\n   paginate; detail per call). Categorize (use the agent's own `call_category` data-collection if\n   present), count tool fires, and note failure modes. This tells you which paths matter and at what\n   frequency.\n\nIf you find duplicated logic anywhere (repeated boilerplate across nodes, per-node scoping the base\nalready has, scaffolding duplicating a system tool), simply remove it in the simplified build —\nafter confirming it is a true duplicate, not a variant.\n\n## Phase 2 — Build the ground-truth test suite BEFORE simplifying\n\nDerive test scenarios from **both** sources:\n- **The complex agent itself**: every distinct behaviour in the node `additional_prompt`s, the base\n  prompt's flows (e.g. an operator-transfer escalation sequence), each qualifying-transfer case,\n  say-node wording, dispatch-tool paths, procedure flows, tool-consent rules, deprecated services,\n  coverage/country rules, language policy.\n- **Real conversations**: one scenario per distinct path/category actually observed, phrased as a\n  simulated-user persona (reference the source `conv_` id in the scenario for traceability).\n\nTwo complementary instruments — pick per behaviour:\n\n| | `simulation` tests (multi-turn) | `llm` tests (single-turn) |\n|---|---|---|\n| What | Simulated user persona converses N turns; whole conversation judged against `success_conditions` | Fixed `chat_history`; judge the agent's NEXT reply against `success_condition` |\n| Use for | Procedure-driven and multi-step flows (diagnostics, escalation sequences, retention) | Sharp single decision points (don't transfer on first demand; reply language; refusal) |\n| Avoid for | — | Any flow whose first agent turn is a tool call (`start_procedure`, `language_detection`): the judge sees no text → `unknown`. Use a simulation test instead. |\n\nCriteria-writing rules (learned the hard way):\n- **Verify every fact in a criterion against the KB before asserting it.** Do not label something a\n  hallucination until you've grepped the KB (a \"made-up\" USSD code turned out to be documented).\n- **Accept reasonable behaviour.** If answering a direct question briefly before a still-needed\n  transfer is fine, say so in the criterion. Over-strict criteria create false failures you'll then\n  \"fix\" wrongly.\n- **Don't let the scenario undermine the premise.** If testing \"plain SIM fails where partner is\n  4G-only\", pin a country that actually has no 3G (check the KB), or the agent will be correct and\n  the test wrong.\n- **Don't test `end_call`.** Test harnesses don't simulate call termination reliably, and the caller\n  hangs up anyway. Testing that the agent does NOT hang up prematurely is fine.\n- One criterion per saved test where possible (clean per-behaviour pass/fail in the dashboard).\n\nMechanics:\n- Save as **native tests** (`POST /v1/convai/agent-testing/create`) and attach to both the\n  **original (reference) agent** and the **simplified copy** via\n  `platform_settings.testing.attached_tests` — dashboard-visible, re-runnable by the client, and\n  the same test objects score both sides of the A/B.\n- Mock every webhook/client tool with **`tool_mock_overrides`** (per-test inline mock bodies keyed\n  by tool id: `{tool_id: [{\"mock_result\": \"<json string>\"}]}` with\n  `tool_mock_config: {mocking_strategy: \"all\", ...}`). Nothing hits real endpoints and no global\n  tool object is edited.\n- Seed `dynamic_variables` from a real call (scrubbed), and include `system__call_sid`,\n  `system__caller_id`, `system__conversation_id`, `system__called_number` — simulations error with\n  `missing_dynamic_variables` otherwise if any tool references them.\n- Tests already attached to the original agent are shared objects: adding new tests alongside them\n  is fine, but never edit or remove the existing ones.\n\n**Baseline the original agent on the suite first.** Its failures are improvement opportunities, not\nexcuses (\"equivalently bad\" is not the goal).\n\n## Phase 3 — Simplify\n\nDefault target: **single-agent (1-node workflow)** + base prompt + focused rules + procedures + KB.\nKeep the workflow only where determinism is genuinely required and prompt enforcement proves\nunreliable in tests.\n\n- Flatten the workflow (`PATCH` with a start-node-only workflow). Check what the main/hub node's\n  `additional_prompt` actually adds beyond the base prompt — if it adds nothing (or duplicates it),\n  there is nothing to port from that node.\n- Port each node's **unique** logic (not any repeated boilerplate) into either:\n  - **focused prompt rules** — short `#`-titled sections, one behaviour each; or\n  - **procedures** — for genuinely multi-step flows (diagnostics, retention). See the procedures\n    API lifecycle below; procedures on the copy are safe to edit (per-agent clones).\n- Deterministic phrasing the client mandated (say nodes) → encode single verbatim phrases as\n  explicit prompt rules and test them. **Mandated multi-step sequences** (e.g. \"clarify twice, ask\n  a feedback question, only then transfer\") belong in a **procedure** — step ordering is exactly\n  what procedures are for, and prompt-rule enforcement of step order is less stable (the model\n  tends to jump to the terminal action once the customer has pushed enough times, skipping the\n  last mandated step). Whichever mechanism you use, test the full sequence multi-turn.\n- Common config fixes while you're there: `PATCH` rejects a body containing both resolved `tools`\n  and `tool_ids` — pop `tools`, keep `tool_ids`. Read-modify-write the full `conversation_config`\n  so nothing is wiped.\n\n## Phase 4 — Iterate on failures (the discipline that makes this work)\n\nRun the suite on the copy; then, for **every** non-passing test (on the copy **or** the original):\n1. Read the transcript/rationale once. **Do not re-run to confirm a failure** — one failure means\n   there's something to improve. (Re-running to verify a FIX is fine.)\n2. Classify: **agent bug** (fix with a focused rule/procedure edit, then verify) vs **test bug**\n   (wrong premise, over-strict criterion, wrong instrument — fix the test, never by loosening it\n   past what's actually correct) vs **ungradable artifact** (tool-only turn in an `llm` test →\n   convert to a simulation test or retire the duplicate).\n3. Failures shared by the original agent are still yours to fix — the tests encode ground truth\n   from real calls, and the goal is a better agent, not parity with its defects.\n4. Watch for **rule dilution**: as prompt rules accumulate, earlier enforcement can weaken. If a\n   previously-passing mandated behaviour regresses, strengthen that rule's self-check rather than\n   adding more prose.\n\n## Phase 5 — Final A/B on branches of the original agent\n\nThe cleanest comparison is two branches of the SAME agent:\n- `POST /v1/convai/agents/{id}/branches` with `{parent_version_id, name, description,\n  conversation_config, platform_settings, workflow}` → a branch (e.g. \"Pete\") carrying the\n  simplified config; Main stays untouched. `parent_version_id` = the agent's current `version_id`.\n- `GET /v1/convai/agents/{id}?branch_id=...` to verify each branch's config.\n- Run the full suite per branch: `POST /v1/convai/agents/{id}/run-tests` with\n  `{\"tests\":[{\"test_id\":...}], \"branch_id\": ...}`; poll\n  `GET /v1/convai/test-invocations/{invocation_id}` until all runs are terminal; result is\n  `condition_result.result` (`success` / `failure` / `unknown`; `rationale` may be a dict with\n  `summary`).\n- Report per-branch totals + per-test diffs, with invocation ids (client-verifiable in the\n  dashboard).\n\n## Phase 6 — Report honestly\n\n- State the claim precisely: \"reproduces every tested behaviour of the original on a much simpler\n  architecture and scores X vs Y on the suite\" — not \"a 1-node agent is inherently better\".\n- Name the caveats: LLM-judged simulations (strong signal, not live traffic), single-run noise,\n  you authored the tests you're scoring against, per-turn token cost of a bigger always-on prompt,\n  and workflow determinism vs prompt enforcement for verbatim phrases.\n- Recommend human review of a few calls and a gradual rollout before any full switch.\n- If you got something wrong along the way, correct the record explicitly in the report.\n\n## API gotchas appendix\n\n**Procedures (draft → commit lifecycle — everything is branch-scoped):**\n- `POST /v1/convai/agents/{id}/branches/{br}/procedures` `{name, content, type: \"free_form\",\n  trigger?}` → returns `procedure_id`, but creates the procedure **as a draft** and adds its ref to\n  your user's agent draft. It will NOT appear in the committed agent/branch procedure lists yet —\n  do not conclude the API \"didn't attach\" (checking the committed view after create is the classic\n  mistake).\n- `GET .../procedures/status` → the draft refs; `GET agent?include_draft=true` → draft config.\n- `POST .../procedures/compile` → **preview only**: validates and returns the compiled workflow +\n  errors. It commits nothing.\n- **Commit = `PATCH /v1/convai/agents/{id}?branch_id={br}`** (an empty JSON body works): it\n  resolves procedure refs from your draft into the new committed version, then deletes the drafts.\n  After this, the procedure appears in `GET agent` `.procedures`, the branch list, and direct GET.\n- Edit content: `PATCH .../procedures/{pid}/draft` (full body: name, content, type, trigger) →\n  commit PATCH. Delete: `DELETE .../procedures/{pid}` (removes from draft set) → commit PATCH.\n- Procedure `content` = markdown with `---` frontmatter (`name`, `trigger`) + steps. Reference KB\n  docs as `[kb id=\"...\" name=\"...\"]`, tools as `[tool id=\"...\" name=\"...\"]`, system tools as\n  `[system_tool id=\"transfer_to_number\" name=\"...\"]`.\n- Procedures are per-agent (cloned when an agent is duplicated) — but the KB/tools they reference\n  are global.\n\n**Testing:**\n- Saved-test update is `PUT` (full replace — a partial body silently wipes omitted fields), not\n  `PATCH` (405).\n- `simulate-conversation` (ad-hoc, no saved artifact) runs the LIVE agent config — to A/B a\n  candidate config you must actually PATCH it (or use branches + `run-tests`).\n- A test whose judged turn is a bare tool call grades `unknown` — instrument mismatch, not failure.\n- Multi-turn sims have some single-run flakiness. Don't paper over it with repeat runs — testing\n  has cost; prefer broadening the suite with more varied scenarios, and investigate any failure\n  rather than re-rolling it.\n\n**Agent config:**\n- A language preset makes a language switchable by `language_detection`; but if a `{{language}}`\n  dynamic variable anchors the default, add an explicit \"the spoken language wins\" prompt rule.\n- Conversation-detail endpoint rate-limits hard — keep concurrency ~4-5 with retries.\n"
}

SHA-256: 606c6bb7799302eb3473473c662e0321bdcf975fb16abf81079da82e8cbb122f