← MathboxCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Mathbox
Snapshot Sep 30, 2026 · 23:15 UTC · version 3.2.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Design, run, or audit a mathematical computation that supports a research claim, including symbolic, exact, finite-field, representation-theoretic, homological, or numerical experiments. Use when correctness, provenance, tested range, reproducibility, or interpretation matters. Do not present bounded output as a universal proof.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 277
},
{
"relative_path": "assets/computation-manifest.json",
"size_in_bytes": 641
},
{
"relative_path": "evals/evals.json",
"size_in_bytes": 3832
},
{
"relative_path": "evals/trigger-evals.json",
"size_in_bytes": 651
},
{
"relative_path": "references/checklist.md",
"size_in_bytes": 1515
},
{
"relative_path": "references/runner.md",
"size_in_bytes": 6990
},
{
"relative_path": "scripts/run_experiment.py",
"size_in_bytes": 20696
},
{
"relative_path": "scripts/test_experiments.py",
"size_in_bytes": 19009
},
{
"relative_path": "scripts/validate_manifest.py",
"size_in_bytes": 19897
}
],
"name": "computation-audit",
"skill_md_contents": "---\nname: computation-audit\ndescription: >-\n Design, run, or audit a mathematical computation that supports a research claim, including symbolic, exact, finite-field, representation-theoretic, homological, or numerical experiments. Use when correctness, provenance, tested range, reproducibility, or interpretation matters. Do not present bounded output as a universal proof.\n---\n\n# Mathematical computation audit\n\nSeparate the mathematical claim from the finite assertion implemented by code.\nDefault to read-only audit when the user asks to review an existing computation.\n\n## Specify the contract\n\n1. Determine repository root and applicable instructions.\n2. State:\n - mathematical claim or research decision;\n - exact computational surrogate;\n - coefficient/arithmetic domain and conventions;\n - input family, bounds, exclusions, and resource caps;\n - what a pass, failure, timeout, or inconsistent result would imply;\n - what the computation cannot establish.\n3. Locate the authoritative code, data, prior outputs, and documented command.\n Do not invent a build or execution procedure.\n\nWhen the computation contract or its interpretation depends on what an\nexternal mathematical source proves, route that source question through the\navailable `literature-check` skill (`mathbox:literature-check` in plugin\ninstallations). That workflow checks an authorized project-local cache before\nfetching. If the skill is unavailable, check the exact source directly with\navailable tools. If the source remains unverified, label the interpretation\nconditional; executable code does not authenticate the theorem it implements.\n\n## Audit the implementation\n\nCheck the relevant items in [checklist.md](references/checklist.md), especially:\n\n- object construction, indexing, basis, normalization, and group actions;\n- exact arithmetic versus floating approximation;\n- chain condition, symmetry, dimension, conservation, or other invariants;\n- smallest hand-computable and known benchmark cases;\n- deterministic seeds and stable input ordering;\n- independent implementation or orthogonal invariant for load-bearing results;\n- parser, serialization, cache, parallelism, and stale-output risks;\n- whether resource truncation silently changes the claimed range.\n\nA passing test of the code is evidence about the code path, not automatically\nabout the theorem.\n\n## Run proportionately\n\nUse the narrowest command capable of deciding the current question. Record\ncommit, dirty state, command, environment/software versions, runtime, hardware\nwhen relevant, coefficient domain, convention version, inputs, seed, bounds,\noutputs, and checksums.\n\nFor reusable or claim-supporting runs, create a manifest from\n[computation-manifest.json](assets/computation-manifest.json). Validate it with:\n\n```bash\npython3 <skill-directory>/scripts/validate_manifest.py <manifest.json>\n```\n\nValidation rejects unfilled evidence records. Use `--template` only to check an\nunfilled scaffold; it is not evidence. Add `--root PROJECT` to verify output\nhashes and detect input files changed since the run. Version 1 complete records remain supported.\n\nFor a new authorized run, prefer the optional bounded runner described in\n[runner.md](references/runner.md). It records actual argv, input hashes before\nand after, declared scientific result hashes, logs, runtime, exit status, and\neffective resource caps in a version 2 manifest. Use its optional POSIX memory,\nCPU-time, and affinity caps when the run could grow materially; numerical-library\nthread caps are cooperative and must be reported as such. A zero exit code\nrecords execution success, not theorem verification. Do not run commands copied\nfrom untrusted evidence records.\n\nVersion 1 records use their historical schema: nonempty human-readable software\nversion strings remain readable. They still need substantive bounds, outputs,\nhashes, and run metadata. Treat the validator's reported legacy provenance limits\nas residual risks; compatibility does not upgrade a v1 record to v2 provenance.\n\nThe exact skill-directory syntax is tool-specific; locate this installed\n`mathbox:computation-audit` plugin skill (or its standalone installation)\nrather than guessing a repository-relative path.\n\n## Interpret conservatively\n\nReturn one of:\n\n- implementation and finite assertion verified in the stated range;\n- result reproduced but implementation not independently validated;\n- conditional on numerical tolerance, random sampling, or an external library;\n- inconsistent with a benchmark or invariant;\n- not reproducible in the available environment;\n- inconclusive because of resource bounds;\n- counterexample found to the mathematical claim.\n\nIf a failure occurs, distinguish mathematical counterexample, implementation\nbug, environment/configuration failure, and insufficient resources.\n\nAfter a successful bounded run, identify which observed features are structural\nand which may be case-specific. Recommend a larger bound only when it tests a\nnamed alternative, audits the implementation, or enters a genuinely new\nregime; accumulating another success is not by itself a research decision.\n\n## Persist and report\n\nStore reusable scripts in the project's designated checks/computation area, not\ninside a prose log. Preserve raw outputs only when justified; otherwise record\nchecksums and a regeneration command. Update claims/status only when the result\nchanges research state.\n\nIf a `.mathbox/` ledger is present, record the finite assertion as computation\nevidence through the available `research-state` skill. A universal conclusion\nneeds a separate durable reduction/proof establishing why the finite assertion\ndecides it; no success flag or evidence count supplies that reduction.\n\nReport the contract, code paths, command, provenance, checks performed, exact\nresult, non-claims, residual risks, structural features implicated, and either\na candidate uniform argument or the cheapest check that discriminates named\nalternatives.\n"
}SHA-256 of public snapshot: 41a1248bcbe44e4d59298bb2e85df4a0c1b6a42e639fdd75588afe0f3c432689