← VapiCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Vapi
Snapshot Sep 30, 2026 · 23:15 UTC · version 1.2.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "simulations",
"description": "Design, create, run, monitor, and maintain Vapi Simulations for assistants and squads. Use for simulation personalities, scenarios, structured-output success criteria, simulations, suites, chat or voice runs, tool mocks, target variables, lifecycle webhooks, regression coverage, CI quality gates, run-result analysis, and simulation API validation errors. Do not use for fixed-turn mock-conversation Evals unless the user is deciding between Evals and Simulations.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 192
},
{
"relative_path": "references/api-reference.md",
"size_in_bytes": 11626
}
],
"skill_md_contents": "---\nname: simulations\ndescription: Design, create, run, monitor, and maintain Vapi Simulations for assistants and squads. Use for simulation personalities, scenarios, structured-output success criteria, simulations, suites, chat or voice runs, tool mocks, target variables, lifecycle webhooks, regression coverage, CI quality gates, run-result analysis, and simulation API validation errors. Do not use for fixed-turn mock-conversation Evals unless the user is deciding between Evals and Simulations.\nlicense: MIT\n---\n\n# Vapi Simulations\n\nBuild realistic conversation tests in five layers: a personality controls the AI tester, a scenario defines its intent and measurable outcomes, a simulation pairs them, a suite groups simulations, and a run executes them against an assistant or squad.\n\n## Source and Safety Rules\n\n- Verify live payloads against the current Vapi documentation MCP, API reference, or public OpenAPI before sending them. Simulations use the `/eval/simulation` API family.\n- Never print, request in chat, or embed API keys, provider secrets, credential values, private webhook URLs, or real customer data.\n- Treat running a simulation as an external action. It can consume credits, use concurrency, send webhooks, and call the target's real tools unless they are mocked.\n- Do not run, cancel, update, or delete resources unless the user clearly requests that operation. Draft configurations when mutation is not requested.\n- Resolve every assistant, squad, personality, scenario, simulation, suite, tool, structured-output, and credential ID from user input or the API. Never invent an ID.\n- Do not create legacy Test Suites. Use Evals for deterministic turn-by-turn checks and Simulations for dynamic conversations over chat or voice.\n\n## Procedure\n\n1. Choose the test type and execution mode.\n - Use Simulations for multi-turn behavior, personality variation, squad handoffs, realistic tool paths, or audio behavior.\n - Use Evals instead when the requirement is an exact response, regex, fixed mock conversation, or precise tool-call argument check.\n - Return a test plan or payload when the user asks to design, draft, review, or explain. Perform live mutations only when explicitly requested and `VAPI_API_KEY` is available.\n\n2. Inspect the target and existing test resources.\n - Fetch the assistant or squad and identify its core paths, guardrails, tools, variables, languages, and failure behavior.\n - List existing personalities, scenarios, simulations, suites, and reusable structured outputs before creating duplicates.\n - Reuse an existing resource only when its intent and configuration match unambiguously. Otherwise create a clearly named new resource or ask the user to choose among plausible matches.\n\n3. Design coverage before payloads.\n - Start with one smoke simulation for the core path, one or two required Boolean outcomes, chat transport, and one iteration.\n - Add regression simulations for repaired defects. Add separate edge cases for ambiguity, interruption, refusal, unavailable dependencies, failed tools, escalation, and handoffs.\n - Keep scenario intent, personality behavior, and evaluation criteria independent so each can be reused.\n - Name resources by behavior and expected outcome, not implementation details.\n\n4. Define the personality.\n - Prefer a suitable existing personality when available.\n - When creating one, provide a complete valid assistant configuration for the AI tester. Put stable temperament, speaking style, and caller behavior in its system prompt; put the situation-specific goal in the scenario.\n - Use the `create-assistant` skill to assemble or validate the personality's assistant configuration when available.\n - Configure voice and transcriber only when voice runs need them. Chat runs use the personality's model but skip its audio path.\n\n5. Define the scenario and evaluations.\n - Write `instructions` as the AI tester's intent and facts. Describe the goal and constraints without scripting the target assistant's answer.\n - Make each evaluation measure one observable outcome. Prefer descriptive Boolean outputs for pass/fail facts and numeric outputs for thresholds.\n - Provide either `structuredOutputId` or inline `structuredOutput`, never both. Inline outputs require `name` and a JSON `schema`.\n - Match the expected `value` type to the evaluated primitive. Use `=` or `!=` for Boolean and string; numeric types also support `>`, `<`, `>=`, and `<=`.\n - Keep important criteria `required: true`. Use optional criteria only for diagnostics that must not fail the simulation.\n - Object structured outputs may be evaluated through a primitive leaf using `path`. Do not compare an object or array directly.\n\n6. Isolate side effects and runtime context.\n - Inspect the target's configured tools before every run. Mock any tool whose real execution could write data, contact people, spend money, or make the test non-deterministic.\n - Match each `toolMocks[].toolName` exactly. The mock `result` is always a string; encode JSON as a string when the target expects JSON-shaped output.\n - Assume every unmocked tool remains live in both chat and voice simulations.\n - Put test values for `{{variables}}` in `targetOverrides.variableValues`. Use synthetic data and keep secrets in Vapi credentials.\n - Configure `simulation.run.started` or `simulation.run.ended` hooks only when requested. Prefer `server.credentialId` to inline authorization headers.\n\n7. Create and verify reusable resources.\n - Create in dependency order: personality and scenario, then simulation, then optional suite.\n - Require `201` for create operations. Verify returned IDs and the fields that define the test.\n - For updates, fetch the current resource first. Omit unrelated scalar fields and send the complete intended value for any array being changed; suite `simulationIds` and `targetAssignments` replace their existing arrays.\n - Re-fetch after update. Deleting a suite or other simulation resource is permanent; verify the exact ID and dependency impact first.\n\n8. Run deliberately.\n - Prefer `vapi.webchat` for fast prompt, tool, and conversation-logic iteration.\n - Use `vapi.websocket` for speech recognition, voice output, interruptions, recordings, or final end-to-end validation.\n - Start with one iteration. Increase iterations only to measure behavioral consistency after a single run is valid.\n - Before sending the run, recap the target, simulations or suite, transport, iterations, tool mocks, and any remaining live side effects.\n - Create the run with `POST /eval/simulation/run` and require `201`. Return the run ID and dashboard `url` when present.\n\n9. Monitor and diagnose results.\n - Poll `GET /eval/simulation/run/{id}` until `status` is `ended`; do not treat `queued` or `running` as success.\n - Fetch `GET /eval/simulation/run/{id}/item` and inspect every item. A passing group has items to evaluate, zero failed or canceled items, and every required evaluation passes.\n - Report actual versus expected values, extraction errors, skipped evaluations, failure reasons, transcript evidence, transport, and iteration number.\n - Diagnose the failing layer before changing the assistant: target runtime failure, scenario ambiguity, personality behavior, tool mock mismatch, structured-output extraction, or genuine assistant behavior.\n - Keep the evaluation stable when fixing the assistant. Change expected criteria only when the business requirement changed.\n\n10. Handle failures honestly.\n - On `400`, compare the request with the current schema and correct one unambiguous validation issue before at most one retry.\n - On `401` or `403`, stop for authentication or permission. On `404`, report the missing dependency. On `409` or concurrency errors, inspect `GET /eval/simulation/concurrency` and active runs. On `5xx`, report the service failure.\n - Cancel only queued or running groups or items. Never claim a run, cancellation, mutation, or pass succeeded until the corresponding API response is verified.\n\n## API Implementation\n\nRead [Simulation API Reference](references/api-reference.md) before producing REST code, making a live request, configuring hooks or mocks, or interpreting run results. Use direct REST unless the current official Vapi SDK documentation explicitly exposes the required simulation resource and method; never invent SDK method names.\n\n## Output Contract\n\nReturn only the sections relevant to the request:\n\n- Test strategy: target behavior, coverage, and why Simulation rather than Eval\n- Resource plan: personality, scenario, evaluations, simulation, and suite\n- Side-effect review: mocked tools, live tools, hooks, variables, transport, iterations, and expected cost/concurrency impact\n- Save-ready JSON or implementation code\n- Created resource IDs and verified fields, when mutations succeeded\n- Run ID, dashboard URL, status, item counts, and per-evaluation evidence, when a run was requested\n- Failure diagnosis and the smallest recommended next change\n\n## Public Sources\n\n- [Simulations overview](https://docs.vapi.ai/observability/simulations-overview)\n- [Simulations quickstart](https://docs.vapi.ai/observability/simulations-quickstart)\n- [Simulations advanced](https://docs.vapi.ai/observability/simulations-advanced)\n- [Manage simulations](https://docs.vapi.ai/observability/simulations-manage)\n- [Vapi API reference index and OpenAPI](https://docs.vapi.ai/llms.txt)\n"
}SHA-256: eee0349449cb13d1ff57d454b9a68e4d85baf7b6c44123e8c744293d5a6e9135