{"id":20807,"plugin_id":"plugins_6aa91bf89030819199fe5546d155f32f","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:33.197Z","digest":"33382d38c4501087582d95668d58c355e70bde3c9c5f8445f0290a5e9b7760cf","against":null,"payload":{"description":"Compare an AI-produced artifact with the accepted final artifact and quantify the human correction and review it required. Use when someone asks whether AI saved work or shifted effort into editing, review, rework, or escalation. Do not infer time saved, authorship, or causality from a text diff alone.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":239},{"relative_path":"evals/cases.json","size_in_bytes":876},{"relative_path":"references/classification.md","size_in_bytes":1218},{"relative_path":"scripts/compare_artifacts.py","size_in_bytes":3448}],"name":"human-review-tax","skill_md_contents":"---\nname: human-review-tax\ndescription: Compare an AI-produced artifact with the accepted final artifact and quantify the human correction and review it required. Use when someone asks whether AI saved work or shifted effort into editing, review, rework, or escalation. Do not infer time saved, authorship, or causality from a text diff alone.\nlicense: MIT\n---\n\n# Human Review Tax\n\nMeasure the work between AI output and accepted output. A large diff is not automatically bad, and a small diff is not automatically correct.\n\n## Boundary\n\n- Analyze artifacts and workflow evidence, not individual reviewer performance.\n- Do not infer effort, quality, authorship, savings or causality from diff size.\n- Preserve both source artifacts; write derived reports or diffs to new files only.\n\n## Required inputs\n\n- The first AI-produced artifact or identifiable AI-assisted revision.\n- The accepted final artifact, or the latest available revision clearly labeled as not yet accepted.\n- The task and quality criteria used by the reviewer.\n\nOptional evidence includes timestamps, comments, tracked changes, review rounds, ticket history, test failures, and the reviewer's own time estimate.\n\nIf the first AI artifact cannot be isolated, analyze the review process qualitatively and label attribution unknown.\n\n## Workflow\n\n1. Establish comparable versions. Record hashes, timestamps, formats, and acceptance status. Do not compare unrelated drafts merely because their filenames match.\n2. Run `scripts/compare_artifacts.py` for a deterministic structural baseline. Its counts describe changed material, not effort or quality.\n3. Review every material change and classify its primary reason: correctness, omission, unsupported claim, task misunderstanding, judgment, structure, style, policy or risk, formatting, or scope added after the AI draft.\n4. Separate repair from preference. A reviewer choosing another valid style is not the same as fixing wrong work.\n5. Reconstruct the review loop from available evidence: review rounds, elapsed time, active human time, blockers, reopens, and escalations. Keep reported time separate from inferred time.\n6. Identify displaced work. Name the upstream effort AI reduced and the downstream work it created.\n7. Recommend at most three changes tied to recurrent evidence: improve the task contract, constrain the AI step, add a verifier, change the handoff, or stop using AI for that part.\n\nRead [references/classification.md](references/classification.md) before classifying changes in a high-stakes artifact.\n\n## Completion criterion\n\nEvery material difference is classified or marked unresolved; repair is separated from preference and changed scope; time values identify their source; accepted status is explicit; and no statement of savings or causality exceeds the available evidence.\n\n## Output\n\n1. Net assessment in plain language\n2. Artifact identity and comparison coverage\n3. Review ledger by change category\n4. Review loop: rounds, elapsed time, active time, and evidence source\n5. Work removed versus work created\n6. Repeated failure pattern, if the evidence supports one\n7. Up to three workflow changes\n8. Evidence limits\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}