← ElevenLabsCONTENT HISTORY

Update to ElevenLabs

Snapshot Sep 30, 2026 · 22:53 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "fix-agent-qa-ticket",
  "description": "Fix an Architect QA triage ticket (agtqa_*) end-to-end via the raw agents API: fetch the ticket + conversation, find the root-cause procedure or prompt text, fix it on a branch, add + run a simulation test proving the fix, then post a summary comment on the ticket. Use when asked to \"fix agtqa_...\", \"address this QA ticket\", \"triage and fix this Architect ticket\", or given an agtqa_* id/URL.\n",
  "included_files": [],
  "skill_md_contents": "---\nname: fix-agent-qa-ticket\ndescription: >\n  Fix an Architect QA triage ticket (agtqa_*) end-to-end via the raw agents\n  API: fetch the ticket + conversation, find the root-cause procedure or prompt\n  text, fix it on a branch, add + run a simulation test proving the fix, then post\n  a summary comment on the ticket. Use when asked to \"fix agtqa_...\", \"address this\n  QA ticket\", \"triage and fix this Architect ticket\", or given an agtqa_* id/URL.\n---\n\n# Fix a QA ticket (agtqa_*)\n\n`agtqa_*` = an Agent Conversation Triage Ticket: a reviewer's flagged comments\nagainst specific turns of a real Architect/ConvAI conversation. This skill turns\none of those tickets into a verified, branch-scoped fix, entirely through the raw\nConvAI REST API (host `https://api.elevenlabs.io`, header `xi-api-key`). It never\ntouches main directly — it stages everything on a branch and leaves merging to\nthe user.\n\n**Needs from the user**: an API key and a ticket id (`agtqa_...`) or a URL\ncontaining one. Everything else — agent id, conversation id, branch — is derived\nfrom the ticket.\n\n**Platform-only.** Do not write repo code as part of this skill unless the root\ncause is actually a code bug (rare — most triage findings are prompt/procedure\ngaps). If code IS the root cause, say so and hand off instead of guessing at a\nplatform-side workaround.\n\n## 1. Fetch the ticket\n\n```bash\ncurl -s \"https://api.elevenlabs.io/v1/convai/conversation-triage-tickets/$TICKET_ID\" \\\n  -H \"xi-api-key: $API_KEY\"\n```\n\nReturns `conversation_id`, `agent_id`, `qa_comment`, `ticket_comments[]`,\n`turn_comments[]` (each with `turn_index` + `comment`), `status`\n(`open`/`in_progress`/`resolved`).\n\n## 2. Fetch the conversation and read the flagged turns\n\n```bash\ncurl -s \"https://api.elevenlabs.io/v1/convai/conversations/$CONVERSATION_ID\" \\\n  -H \"xi-api-key: $API_KEY\"\n```\n\n`transcript[]` has per-turn `role`, `message`, `tool_calls` (each with\n`tool_name` + `params_as_json`). Read a window around every `turn_index` named in\n`turn_comments` — usually a couple turns before and after — to understand what\nthe agent actually did and said. The reviewer's comments tell you *what's wrong*;\nthe transcript tells you *why*. Don't skip straight to guessing the fix — find\nthe specific tool call, prompt gap, or missing instruction that produced the bad\nturn.\n\nThe conversation also carries `branch_id` and `version_id` — useful context, but\n**do not build your fix branch off this conversation's branch** unless it's\ndemonstrably the right base (check `is_archived` / whether it's actually merged\ninto main — see step 3). It's just where the flagged conversation happened to run.\n\n## 3. Create a fix branch off main's actual tip\n\nGet the agent and find its real main branch:\n\n```bash\ncurl -s \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID\" -H \"xi-api-key: $API_KEY\"\n# -> .main_branch_id, .branch_id (top-level is usually main)\n```\n\nThen fetch that branch to get its true HEAD version (do not assume — a stray\npersonal/archived branch can look tempting but have unrelated commits in its\nancestry):\n\n```bash\ncurl -s \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$MAIN_BRANCH_ID\" \\\n  -H \"xi-api-key: $API_KEY\"\n# -> .most_recent_versions[0].id  is main's real tip version_id\n```\n\nCreate the branch from that exact version:\n\n```bash\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"name\": \"<you>/qa-<ticket-suffix>-<short-desc>\", \"description\": \"Fix for '\"$TICKET_ID\"'\", \"parent_version_id\": \"<main-tip-version-id>\"}'\n```\n\nRequired fields are `name`, `description`, `parent_version_id` (all three, or\nyou get a 422 listing what's missing). Response: `{created_branch_id,\ncreated_version_id}`.\n\n## 4. Find and fix the root cause\n\nUsually one of:\n- **A procedure** missing an instruction (list via\n  `GET .../agents/{id}/branches/{b}/procedures`, fetch full content via\n  `GET .../procedures/{pid}`). Search procedure names/triggers for the relevant\n  topic (dashboards, refunds, whatever the ticket concerns).\n- **The system prompt** (`conversation_config.agent.prompt.prompt` off\n  `GET /v1/convai/agents/{id}?branch_id={b}`).\n\n### Editing an existing procedure (draft + publish dance)\n\nEditing an existing procedure is **two steps**, not one — there is no direct\ncommit-on-PATCH for procedures:\n\n```bash\n# 1. Stage the draft (full content: frontmatter + body, or per-field if the\n#    procedure isn't markdown-frontmatter style)\ncurl -s -X PATCH \\\n  \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/branches/$BRANCH_ID/procedures/$PROCEDURE_ID/draft\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"name\": \"...\", \"content\": \"...\", \"type\": \"free_form\", \"trigger\": \"...\"}'\n\n# 2. Publish ALL pending procedure drafts on the branch by re-committing the\n#    agent with any partial merge body (re-sending the branch's current,\n#    unmodified prompt is the simplest no-op payload). Fetch the CURRENT prompt\n#    scoped to $BRANCH_ID first (not main's) so you don't clobber other\n#    branch-local differences:\ncurl -s \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID\" -H \"xi-api-key: $API_KEY\" \\\n  | jq -r '.conversation_config.agent.prompt.prompt' > current_prompt.txt\n\ncurl -s -X PATCH \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID?branch_id=$BRANCH_ID&version_description=...\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d \"{\\\"conversation_config\\\": {\\\"agent\\\": {\\\"prompt\\\": {\\\"prompt\\\": $(jq -Rs . < current_prompt.txt)}}}}\"\n```\n\nVerify by re-`GET`ting the procedure and confirming the `version_id` changed and\nthe new content is present.\n\nA direct `PATCH .../procedures/{pid}` (no `/draft` suffix) does not exist —\nreturns 405. Creating a brand-new procedure via `POST .../procedures` commits\nimmediately, but don't use that to \"replace\" an existing one in place — it\nleaves two procedures with overlapping/duplicate triggers, which is worse than\nthe bug you're fixing.\n\n### Editing the prompt / criteria directly\n\nNo draft dance needed — a partial-merge `PATCH /v1/convai/agents/{id}?branch_id={b}`\ncommits straight to branch HEAD and returns a new `version_id`:\n- prompt: `{\"conversation_config\":{\"agent\":{\"prompt\":{\"prompt\":\"...\"}}}}`\n- criteria: `{\"platform_settings\":{\"evaluation\":{\"criteria\":[...]}}}` (send the\n  full array; each `conversation_goal_prompt` max 2000 chars)\n\n## 5. Add a simulation test that reproduces the ticket's scenario\n\n**Always use a `simulation` test for Architect/DOM agents** — `llm`/`response`\ntests only grade the agent's first action, which for Architect is almost always\n`start_procedure`, so the judge returns useless `unknown`/`failure` verdicts.\n\n```bash\n# Optional: group tests for this ticket\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/agent-testing/folders\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"name\": \"QA '\"$TICKET_ID\"'\"}'\n\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/agent-testing/create\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\n    \"type\": \"simulation\",\n    \"name\": \"<short description> ('\"$TICKET_ID\"')\",\n    \"parent_folder_id\": \"<folder id or omit>\",\n    \"simulation_scenario\": \"<persona + exactly what they do, mirroring the ticket>\",\n    \"success_conditions\": [\"<checklist item 1>\", \"<checklist item 2>\", ...],\n    \"simulation_max_turns\": 8,\n    \"tool_mock_config\": {\"mocking_strategy\": \"all\", \"fallback_strategy\": \"raise_error\"},\n    \"chat_history\": [{\"role\": \"user\", \"message\": \"...\", \"time_in_call_secs\": 0}],\n    \"dynamic_variables\": {\"tier\": \"enterprise\", \"product\": \"conversational_ai\", \"objective\": \"...\", \"userInfo\": \"{}\", \"chatHistory\": \"[]\"}\n  }'\n```\n\nGotchas that produce 422s:\n- Every `chat_history` entry needs `time_in_call_secs` (even `0`), or you get a\n  `missing` error pointing at that field.\n- `chat_history` must end with a user message.\n- A `system`-type `tool_result` inside `chat_history` needs `is_error` +\n  `tool_has_been_called` fields.\n- `dynamic_variables` should include whatever the prompt templates on\n  (tier/product/objective/userInfo/chatHistory at minimum) or templating crashes\n  mid-run.\n- `mocking_strategy` is an enum of exactly `all` / `selected` / `none` — there\n  is no fourth value. **`\"all\"` ignores `tool_mock_overrides` entirely** — it\n  mocks every tool call with the generic fallback error, even if you've set a\n  per-tool `mock_result`. If the fix you're testing depends on a specific tool\n  actually returning realistic data (e.g. `list_branches` returning a branch\n  list so the agent can resolve a name to an id), you MUST use\n  `mocking_strategy: \"selected\"` with `mocked_tool_ids: [...]` naming every\n  tool you want mocked — `\"all\"` silently no-ops your overrides and you'll see\n  \"no mock matched\" in the transcript even though the override is saved\n  correctly on the test (verify via `GET /v1/convai/agent-testing/{test_id}`\n  if a mock isn't taking effect — the stored config can look right while the\n  run still fails, which means the strategy, not the override, is wrong).\n- `tool_mock_overrides` shape: `{\"<tool_name>\": [{\"mock_result\":\n  \"<json-string>\", \"parameter_conditions\": [], \"is_error\": false}]}`.\n  `mock_result` is a JSON-encoded **string**, not a nested object — build it\n  with `json.dumps(...)` (or equivalent) before embedding it in the request\n  body. An empty `parameter_conditions` list means \"match unconditionally.\"\n- `mocking_strategy: \"all\"` + `fallback_strategy: \"raise_error\"` (no overrides)\n  makes *every* tool call error mid-conversation (\"technical difficulties\").\n  Fine only for testing something that doesn't depend on a tool succeeding\n  (e.g. a confirmation-message wording fix where the tool call itself is\n  incidental); if the flow needs a tool to actually succeed with realistic\n  data, use `selected` + overrides instead of reaching for `call_real_tool`.\n- DOM-write tools (navigate, set_dashboard_filters, etc.) still may not fully\n  green in the sim harness even with `selected` mocking if you haven't mocked\n  every tool in the chain — treat persistent failures there as behavior\n  documentation, not a bug, once you've confirmed the mocking strategy itself\n  isn't the culprit.\n\nWrite the `success_conditions` directly against the reviewer's `turn_comments` —\neach comment should map to a checklist item the grader can verify. Word each\ncondition to name the CORRECT id/value explicitly (e.g. \"uses id X, not the\nraw name string and not id Y\") rather than just \"doesn't guess\" — a vague\ncondition lets a new, differently-wrong failure mode (e.g. passing the name\nitself as the id) slip through as a pass.\n\n**Editing a sim test = delete + recreate.** There is no in-place edit for a\nsimulation test (`PUT` on a test id is response/llm-only and rejects\n`type: simulation`). If a run reveals your mock config was wrong, `DELETE\n/v1/convai/agent-testing/{test_id}`, fix the payload, and `POST .../create`\nagain — then re-attach and re-run.\n\nAttach it to the agent/branch, then run it:\n\n```bash\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/testing/attach-test\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"test_id\": \"'\"$TEST_ID\"'\", \"branch_id\": \"'\"$BRANCH_ID\"'\"}'\n\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/agents/$AGENT_ID/run-tests\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"tests\": [{\"test_id\": \"'\"$TEST_ID\"'\"}], \"branch_id\": \"'\"$BRANCH_ID\"'\", \"repeat_count\": 1}'\n# -> {id: suite_id, test_runs: [{test_run_id, status: \"pending\", ...}]}\n```\n\n## 6. Poll for the result\n\nSim runs take a few minutes. Poll `GET\n/v1/convai/test-invocations/{suite_id}` until `test_runs[0].status` leaves\n`pending`. Use a background poll loop (Bash `run_in_background` or Monitor), not\na foreground sleep — do not busy-wait in the conversation.\n\nOnce terminal, read `test_runs[0].condition_result.result`\n(`success`/`failure`) and `.rationale.messages` (per-criterion grader\nreasoning) to confirm the fix actually produces the intended behavior — don't\njust check the pass/fail bit, skim the rationale for whether it's testing what\nyou think it's testing.\n\nIf it fails: re-read the rationale against the actual procedure/prompt change,\nadjust, and re-run. Don't loosen `success_conditions` to force a pass unless\nthe condition itself was wrong (too strict/loose) — a passed test the reviewer's\nconcern doesn't check is worse than an honest failure.\n\n## 7. Comment on the ticket\n\n```bash\ncurl -s -X POST \"https://api.elevenlabs.io/v1/convai/conversation-triage-tickets/$TICKET_ID/comments\" \\\n  -H \"xi-api-key: $API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"comment\": \"<root cause>\\n\\n<what changed, branch id, not merged>\\n\\n<test id + PASS/FAIL + key rationale line>\\n\\nWritten by <model>, using Claude Code.\"}'\n```\n\nInclude: the root cause in plain language, the branch id (explicitly \"not\nmerged\" — merging is the user's call), the test id and result, and identify\nyourself per repo convention (`Written by {Model}, using {Harness}`).\n\n**Leave ticket `status` as `open`** (don't PATCH it to `resolved`) unless the\nuser explicitly asks you to close it — you fixed and verified on a branch, but\nthe user still needs to review and merge.\n\n## 8. Report back\n\nTell the user: root cause, branch id + that it's unmerged, test id + result,\nand anything you had to work around (wrong-parent branch mistake, a blocked\ndestructive action, an ambiguous base branch) — these are exactly the kind of\njudgment calls the user should sanity-check before merging.\n"
}

SHA-256: 08a079ba5a898ec614d97259a9b9e73a1fd0808129205f94ddd07d0b0aa48f2d