{"id":20806,"plugin_id":"plugins_6aa91bf89030819199fe5546d155f32f","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:33.039Z","digest":"a6a84cdbc5d96acc66c86ea9deef58adf368b5f79f2c07d0f62e5491d3945373","against":null,"payload":{"description":"Reconcile an AI agent's completion claims with durable evidence in the relevant system of record. Use when someone asks whether agent work actually finished, shipped, sent, changed, or reached production. Do not use a passing local check as proof of an unobserved remote or real-world state.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":245},{"relative_path":"evals/cases.json","size_in_bytes":799},{"relative_path":"references/adapters.md","size_in_bytes":1199},{"relative_path":"scripts/validate_outcome.py","size_in_bytes":2897}],"name":"did-it-land","skill_md_contents":"---\nname: did-it-land\ndescription: Reconcile an AI agent's completion claims with durable evidence in the relevant system of record. Use when someone asks whether agent work actually finished, shipped, sent, changed, or reached production. Do not use a passing local check as proof of an unobserved remote or real-world state.\nlicense: MIT\n---\n\n# Did It Land?\n\nTreat the agent's narration as claims, not evidence. Verify the post-condition in the system that owns the state.\n\n## Boundary\n\n- Verification is read-only by default. Do not create, resend, redeploy, approve, or otherwise mutate state to make a check pass.\n- Do not perform a risky live test without explicit authorization.\n- Keep local, committed, reviewed, merged, deployed and production-observed states distinct.\n\n## Workflow\n\n1. Recover the contract. Identify the original request, target, acceptance criteria, exclusions, and any later user corrections. Mark criteria that were never defined.\n2. Extract completion claims from the agent's messages and artifacts. Split compound statements so each row can receive one verdict.\n3. Name the system of record for each claim. Examples: filesystem, git commit, GitHub check, deployment provider, production URL, sent mailbox, calendar event, CRM record, payment ledger, or issue tracker.\n4. Select an independent readback. Prefer a fresh API read, command, or direct artifact inspection over the same agent's summary. Never create or mutate state merely to verify it.\n5. Capture evidence with timestamp, target identity, source, and the smallest excerpt or result needed. Inspect success envelopes: transport success is not application success.\n6. Assign exactly one verdict from the evidence model below. A chain may contain different verdicts, such as committed = verified, deployed = verified, production behavior = unknown.\n7. State the next smallest check needed for every partial, failed, or unknown claim. Do not quietly perform a risky live test.\n\nRead [references/adapters.md](references/adapters.md) for common systems of record. Validate a JSON result with `scripts/validate_outcome.py`.\n\n## Verdicts\n\n- `verified`: fresh evidence directly establishes the claimed post-condition.\n- `partial`: some, but not all, of the claimed state is established.\n- `failed`: fresh evidence contradicts the claim.\n- `claimed-only`: the agent stated it, but no independent evidence was collected.\n- `unknown`: the required system, authority, or observation is unavailable.\n- `not-applicable`: the recovered contract does not require this claim.\n\nDo not convert `claimed-only` or `unknown` into failure. Do not convert a created artifact into proof that it was accepted, sent, merged, deployed, or used.\n\n## Completion criterion\n\nThe check is complete when every acceptance criterion and every material completion claim appears in the evidence ledger; each has a named system of record, verdict, and evidence or explicit evidence gap; and the overall status is no stronger than its weakest required criterion.\n\n## Output\n\nLead with one sentence: what is verified, what is not, and whether required work remains. Then provide:\n\n1. Contract recovered\n2. Evidence ledger\n3. State chain, when work crosses multiple systems\n4. Missing checks and why they were not run\n5. Overall status: verified complete, partially verified, failed, or not verifiable\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}