← Argovance Skill OSCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Argovance Skill OS
Snapshot Sep 30, 2026 · 23:16 UTC · version 1.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Design and implement a durable test and regression system across unit, integration, contract, end-to-end, browser, responsive, accessibility, performance, security-relevant, migration, and visual verification. Use when the primary task is test architecture, coverage strategy, fixtures, CI quality gates, flaky-test control, or regression infrastructure rather than tests for one small feature.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 290
}
],
"name": "engineer-test-and-regression-systems",
"skill_md_contents": "---\nname: engineer-test-and-regression-systems\ndescription: Design and implement a durable test and regression system across unit, integration, contract, end-to-end, browser, responsive, accessibility, performance, security-relevant, migration, and visual verification. Use when the primary task is test architecture, coverage strategy, fixtures, CI quality gates, flaky-test control, or regression infrastructure rather than tests for one small feature.\n---\n\n# Engineer Test and Regression Systems\n\nBuild a test system that detects material failures without optimizing for raw test count or brittle snapshots.\n\nApply the shared [risk-adaptive assurance model](../../shared/expert-system/risk-adaptive-assurance-model.md). Test the required outcome and protected invariants with the smallest reliable system; do not encode an internal implementation path unless security, regulated behavior, protected architecture or an explicit deterministic contract requires it.\n\n## Frame the assurance problem\n\n1. Inventory architecture, critical user journeys, interfaces, data, environments, supported platforms, risks, incident history, current tests, CI, release gates, and observability.\n2. Map requirements and protected behavior to failure modes, direct acceptance oracles, test layers, environments, evidence, owners, and release consequences.\n3. Define what each test layer proves and explicitly does not prove.\n4. Identify deterministic seams, fixtures, test data, mocks, clocks, randomness, external services, cleanup, isolation, and privacy requirements.\n5. Classify tooling, dependencies, environments, and quality thresholds under the shared [decision-authority model](../../shared/expert-system/decision-authority-model.md). Do not introduce them silently.\n\n## Test architecture\n\nSelect applicable layers:\n\n- unit and property tests for local logic;\n- integration for boundaries and persistence;\n- API, schema, event, and consumer-driven contracts;\n- end-to-end journeys and failure recovery;\n- browser, device, responsive, and localization coverage;\n- accessibility semantics and interaction;\n- performance, capacity, reliability, and resource budgets;\n- security-relevant permissions, abuse, input, and dependency checks;\n- migration, rollback, backup, and restore verification;\n- production smoke and synthetic monitoring;\n- visual regression and human perceptual review.\n\nWhen AI behavior is only one layer of a broader test system, include versioned representative and held-out evaluation datasets, expected outcomes/rubrics, prompt/model/retrieval/tool identities, tool-call trajectories and arguments, safety/adversarial cases, failure taxonomy and regression thresholds. Keep probabilistic evaluation distinct from exact deterministic assertions and preserve evaluator limitations. When behavioral AI evaluation is the primary task, route it to `$evaluate-ai-systems`; this skill owns the durable runners, fixtures, CI integration and cross-system regression infrastructure, not the independent AI verdict.\n\nFor material user-facing behavior, drive the real running application when unit or code-level tests cannot establish the outcome. Navigate, interact, trigger state, inspect output and errors, and exercise recovery. For sensitive adversarial or mutation testing, use a bounded isolated candidate copy when direct review could damage the accepted tree or data.\n\nFor visual regression distinguish strict pixel comparison, perceptual comparison, layout geometry, semantic assertions, and human visual approval. Use each only for the observable it can reliably protect. Apply the shared [perceptual quality gates](../../shared/expert-system/perceptual-quality-gates.md); technical tests cannot self-certify premium visual quality.\n\nFor representation-sensitive work, link each important perceptual observable to the representation proof and real runtime state that can establish it. A controlled offline reference render may serve as an art/material/form oracle, but never self-certifies browser pixel truth. Where applicable, test Source-Master-to-publishing-target lineage, representative views and screen-space states, defined versus user-controlled motion, approved publishing tiers and fallbacks, first-contact stability, resting/offscreen/hidden lifecycle behavior, and absence of drift from an authorized calibration checkpoint. Do not require these layers for ordinary DOM/CSS/image interfaces.\n\n## Reliability and governance\n\n- Define test ownership, naming, location, execution commands, runtime budget, parallelism, retries, quarantine, and failure triage. Record only environment factors that can materially change the measurement.\n- Do not hide flaky tests behind unlimited retries. Measure flake rate, isolate the cause, set an owner and removal deadline, and preserve release risk visibility.\n- Version deterministic fixtures and protect secrets and personal data.\n- Define branch, pull-request, pre-release, post-deploy, and rollback gates with evidence retention.\n- Keep the smallest suite that gives required confidence; delete or replace redundant tests only with authority and proof.\n\n## Verification and output\n\nValidate the test system against known seeded failures or safe mutation cases where feasible. Confirm that failures are actionable, reproducible, and linked to requirements. Report false positives, false negatives, runtime, maintenance burden, and uncovered risk.\n\nReturn the coverage and risk map, test architecture, tooling decisions and approvals, fixtures strategy, commands and environments, CI/release gates, visual assurance model, flaky-test policy, implemented assets, validation evidence, gaps, ownership, and rollout plan.\n"
}SHA-256 of public snapshot: ab2ca95d8abb46d6c03fcf370576055b722008f1f2760810d76291b1aa04d117