← Claus Argos Skill OSCONTENT HISTORY

Update to Claus Argos Skill OS

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.16.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Audit, test, benchmark, repair, and safely evolve one or more agent skills. Use when a user asks whether skills work, trigger correctly, overlap, improve results, remain compatible, need better instructions or resources, or should be regression-tested; also use for skill quality reviews, test suites, before/after comparisons, trigger precision, benchmark reports, and evidence-based skill updates.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 242
    },
    {
      "relative_path": "references/evaluation-protocol.md",
      "size_in_bytes": 1048
    },
    {
      "relative_path": "references/system-evolution-regressions.md",
      "size_in_bytes": 4716
    },
    {
      "relative_path": "scripts/static-skill-audit.py",
      "size_in_bytes": 2122
    }
  ],
  "name": "audit-and-evolve-skills",
  "skill_md_contents": "---\nname: audit-and-evolve-skills\ndescription: Audit, test, benchmark, repair, and safely evolve one or more agent skills. Use when a user asks whether skills work, trigger correctly, overlap, improve results, remain compatible, need better instructions or resources, or should be regression-tested; also use for skill quality reviews, test suites, before/after comparisons, trigger precision, benchmark reports, and evidence-based skill updates.\n---\n\n# Audit and Evolve Skills\n\nMeasure usefulness rather than rewarding length or polish.\n\n## Workflow\n\n1. Establish the target skills, runtime, intended users, permitted edits, and acceptance criteria.\n2. Inventory each skill and run `scripts/static-skill-audit.py SKILL_DIR`.\n3. Read `references/evaluation-protocol.md` and create realistic tests covering clear triggers, ambiguous triggers, non-triggers, normal execution, missing inputs, tool failure, unsafe requests, and output compliance.\n   For shared routing/context/authority changes, use [system-evolution-regressions.md](references/system-evolution-regressions.md). Keep observed responses separate from expected criteria; a checklist or structural PASS is not behavioral verification.\n4. Evaluate separately:\n   - discovery and trigger precision;\n   - instruction correctness and context efficiency;\n   - output usefulness and reproducibility;\n   - tool, permission, privacy, and failure behavior;\n   - measurable uplift over the same task without the skill.\n5. Preserve raw prompts, outputs, environment, model/runtime, dates, and grader criteria. Do not present subjective judgment as a benchmark.\n6. Classify findings as `blocking`, `major`, `minor`, or `optional`.\n7. Propose the smallest repair that addresses the demonstrated failure. Do not add text without evidence that it helps.\n8. Obtain approval before modifying installed or shared skills unless the user explicitly requested repair.\n9. Re-run affected tests and report regressions, improvements, remaining uncertainty, and changed files.\n\n## Evolution rules\n\n- Never self-modify from a single anomalous result.\n- Keep held-out tests separate from examples used to write the skill.\n- Require human approval for new permissions, external actions, destructive behavior, memory retention, or changed scope.\n- Preserve previous versions or a recoverable diff.\n- Reject edits that improve a proxy score while weakening the user's real outcome.\n- For recurring or material failure learning use the existing [harness complexity policy](../../shared/expert-system/harness-complexity-policy.md), not automatic new skills. Shared contracts need one canonical home and explicit affected callers; update and test those callers without spreading duplicate rules.\n\n## Output\n\nReturn scope, inventory, trigger matrix, execution scorecard, security/failure findings, proposed repairs, before/after evidence, regressions, and recommendation: `keep`, `repair`, `split`, `merge`, `retire`, or `needs-more-evidence`.\n"
}

SHA-256 of public snapshot: 12cd3410433762be527838aa83b1203275b6e5d3ca0540e10e81816673e0c591