← ServotabCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Servotab
Snapshot Sep 30, 2026 · 23:15 UTC · version 0.6.4
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Verify code or product claims after changes using fresh, risk-matched evidence. Use before saying a bug is fixed, tests pass, requirements are met, or a branch is ready.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 354
},
{
"relative_path": "assets/icon-400.png",
"size_in_bytes": 1436
},
{
"relative_path": "assets/icon.svg",
"size_in_bytes": 1789
}
],
"name": "verify",
"skill_md_contents": "---\nname: verify\ndescription: \"Verify code or product claims after changes using fresh, risk-matched evidence. Use before saying a bug is fixed, tests pass, requirements are met, or a branch is ready.\"\n---\n\n# Verify\n\nEvidence must support the exact claim. Fresh verification is required after the final relevant change, but verification scope should match risk rather than defaulting blindly to the largest test suite.\n\n## Define the claims\n\nList the claims that matter, such as:\n\n- The original bug no longer reproduces.\n- A new behavior matches acceptance criteria.\n- Targeted tests pass.\n- The module builds or type-checks.\n- No existing behavior in the affected area regressed.\n- A migration is safe.\n- The branch is ready to integrate.\n\nFor each claim, identify the command, inspection, or manual scenario that proves it. For decisive cross-boundary probes, name their entry point, relevant path, and material conditions. Freshness and a shared environment do not make a check that bypasses the failing boundary evidence of that boundary's health. A brief coverage note is enough; do not create a second ledger.\n\n## Evidence budget\n\n- Run a check only when its result supports a named claim or can change the next action.\n- Do not calculate hashes without an identity or integrity decision that will use them.\n- Do not rerun unchanged checks or add a second acceptance loop merely to restate existing proof.\n- Stop when every material claim has proportionate fresh evidence; more commands do not automatically create more confidence.\n\n## Evidence maturity without ceremony\n\nKeep capability and effectiveness claims separate:\n\n- A file, rule, tool, or configured capability proves that it exists, not that the task can reach it.\n- A reachable route proves wiring, not successful use or delivery.\n- A focused exercise or passing test proves current behavior under its observed conditions, not general runtime effectiveness. Keep mechanism repair and whole-user-experience recovery separate; residual symptoms do not automatically invalidate a verified contributing fix.\n- A repair verified in the current task proves repair state. Only a later comparable outcome can support a claim that the workflow improved over time.\n- Missing observation is `Not verified`, not automatically a defect.\n\nThese are claim boundaries, not a required scorecard, ledger, report, or extra review loop.\n\n## Host-boundary evidence ladder\n\nFor a host-specific claim, keep these rungs distinct:\n\n1. Source contract\n2. Process-level behavior test\n3. Built artifact or image identity\n4. Activated runtime identity\n5. Exact named-host surface\n6. Owner-observed behavior\n\nEach rung supports the next investigation step, not the claim above it. Local or dev-browser success is not named-host acceptance; a successful build is not proof that the intended runtime is active; deployment is not owner-observed behavior. After the final relevant change, verify the artifact and activated runtime identities, then obtain fresh acceptance on the exact named host when that is the contract. If the required deployment or owner observation is not authorized or available, mark the higher claim `Not verified`.\n\n## Verification ladder\n\n### Level 1: Focused\n\nUse for local, low-risk changes:\n\n- Regression test\n- Affected test file\n- Component or module check\n- Targeted type-check or lint\n- Focused manual interaction\n\n### Level 2: Adjacent\n\nAdd when the change touches shared code or several consumers:\n\n- Package or feature suite\n- Integration tests around the boundary\n- Build for the affected application\n- Representative platform or browser check\n\n### Level 3: Broad\n\nUse for high-risk or integration-ready changes:\n\n- Full relevant test suite\n- Full build\n- Migration dry run\n- End-to-end path\n- Security or compatibility checks\n- Multiple environments when the risk requires it\n\nDo not run Level 3 merely to make a small change look rigorous. Do not stop at Level 1 when shared state, data, security, or public contracts are involved.\n\n## Run and read\n\nFor every command:\n\n1. Run it after the final relevant change.\n2. Read the complete result needed to assess success.\n3. Check exit status, failure counts, warnings, skipped tests, and environment limitations.\n4. Record what it actually proves.\n5. Do not extrapolate beyond that scope.\n\nA prior run before later edits is stale evidence for the affected behavior.\n\n## Non-test checks\n\nInspect:\n\n- Final diff and scope\n- Untracked or generated files\n- Debug output and temporary instrumentation\n- Secrets or sensitive data\n- Schema and fixture consistency\n- Documentation when public behavior or setup changed\n\nFor UI work, include a real rendered or interaction check when practical. Unit tests alone may not prove layout or input behavior.\n\nFor optional host actions, verify the negative capability path before exposure as well as success. Preserve distinct absent, rejected, cancelled, and policy-denied results when the host distinguishes them; a generic error does not prove correct degradation.\n\n## Regression evidence\n\nFor a bug fix, prefer a reproducer or test that would fail under the old behavior. Revert or mutation proof is useful when safe and efficient, but it is not mandatory when it would destabilize the workspace.\n\n## Check the test criterion and close review findings\n\nBefore relying on a green result, consider a plausible incorrect implementation that this check would reject. This is a check on the existing evidence, not a mandatory mutation-testing stage or an extra reviewer loop. Schema presence, file signatures, compilation, a mocked success path, and expected-output updates can all miss the behavior being claimed. Use the nearest available behavioral check or full parser where that is the contract. Pair disappearing errors with the intended successful outcome so that suppressing work or bypassing the observed path cannot masquerade as repair. For state-preservation claims, prefer structured comparison or a demonstrably stable normalized before-image over scattered substring checks; retain semantic fields and ordering, and exclude only understood volatile fields. Keep static checks as static evidence.\n\nFor timing, ownership, recovery, or optional-host changes, inspect the relevant repeated, interrupted, stale, malformed, denied, or accessibility path. Select from these by the actual changed boundary; this is not an exhaustive test matrix for every task.\n\nResolve material review findings against the exact final revision. A finding may be fixed and checked, rejected with a concrete counterexample, or explicitly deferred under applicable authority. Record its disposition in the existing review or task surface. CI green or a merge does not itself resolve a reviewer-identified failure. Reproduce disputed findings instead of trusting either the reviewer or implementer by title.\n\nFor reusable instructions, distinguish discovery, context delivery, method use, and task outcome. Self-reported loading is supporting evidence only. A static assertion about prompt text cannot establish natural-language activation or improved model behavior.\n\n## Blocked verification\n\nWhen a check cannot run:\n\n- State the exact reason.\n- Separate environmental failure from code failure.\n- Run the strongest available alternative.\n- Narrow the completion claim.\n- Give the command or condition needed to complete verification later.\n\n## Output format\n\nUse three categories:\n\n- **Verified:** claim and evidence\n- **Failed:** actual failure and impact\n- **Not verified:** omitted or blocked checks and why\n\nDo not say “all tests pass” when only targeted tests ran. Say exactly which tests passed.\n"
}SHA-256 of public snapshot: b6e9a28f5c9580508cb8465bb537aef49740512254fcd62d5fc350d1ca8dddef