{"id":20804,"plugin_id":"plugins_6aa91bf89030819199fe5546d155f32f","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:32.597Z","digest":"7e874171685b6a8d4b104d36816336818b3658612dac2191174374cfbe992e03","against":null,"payload":{"description":"Reconstruct a completed, failed, expensive, or confusing agent session to find where its trajectory broke and what one change would prevent recurrence. Use when investigating post-hoc session diagnosis, tool-loop analysis, goal drift, premature completion, or repeated human rescue. Do not use as employee performance scoring.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":253},{"relative_path":"evals/cases.json","size_in_bytes":876},{"relative_path":"references/failure-taxonomy.md","size_in_bytes":1398},{"relative_path":"scripts/session_stats.py","size_in_bytes":2773}],"name":"agent-autopsy","skill_md_contents":"---\nname: agent-autopsy\ndescription: Reconstruct a completed, failed, expensive, or confusing agent session to find where its trajectory broke and what one change would prevent recurrence. Use when investigating post-hoc session diagnosis, tool-loop analysis, goal drift, premature completion, or repeated human rescue. Do not use as employee performance scoring.\nlicense: MIT\n---\n\n# Agent Autopsy\n\nFind the first consequential divergence, not merely the last visible error.\n\n## Boundary\n\n- Diagnose one bounded trajectory and preserve the source session unchanged.\n- Do not score a person, infer intent, or convert tool counts into productivity.\n- Do not implement proposed fixes unless the user separately asks for implementation.\n\n## Workflow\n\n1. Freeze the case. Record the session identifier, time range, original task, expected finish, and final observed state. Preserve the source transcript unchanged.\n2. Run `scripts/session_stats.py` for a content-minimized timeline of tool calls, errors, repeated calls, and elapsed gaps. Treat statistics as leads, not diagnoses.\n3. Reconstruct the trajectory as intent, decision, action, observation, and response. Include user corrections, permission denials, tool failures, context compaction, and handoffs.\n4. Mark divergences: goal drift, assumption without evidence, wrong tool or target, ignored observation, retry without new information, instruction conflict, approval gap, premature completion, or verification at the wrong system boundary.\n5. Identify the earliest point where a different decision would probably have changed the outcome. Distinguish triggering event, contributing conditions, and the final symptom.\n6. Measure waste only from observable proxies: repeated tool calls, discarded artifacts, idle gaps, retries, review rounds, and reported cost. Do not convert tokens or elapsed time directly into productivity.\n7. Propose at most three changes. Prefer one tighter task contract, one deterministic guard or verifier, and one instruction or permission change. Each proposal must cite the failure event it addresses.\n8. Define a replay or fixture that would distinguish improvement from a cosmetically better explanation.\n\nRead [references/failure-taxonomy.md](references/failure-taxonomy.md) when multiple failure classes overlap.\n\n## Completion criterion\n\nThe original task and final state are explicit; the timeline accounts for every consequential branch; the first divergence is separated from the final symptom; each recommendation maps to evidence; and a repeatable check is defined for the highest-priority change.\n\n## Output\n\n1. One-paragraph finding\n2. Case boundary and evidence coverage\n3. Timeline of consequential events\n4. First divergence, contributing conditions, and final symptom\n5. Human interventions and recovery points\n6. Observable waste\n7. Up to three corrective changes\n8. Replay plan and remaining unknowns\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}