{"id":18078,"plugin_id":"plugins_6a84a979b5a081918922e037047514b5","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:36.492Z","digest":"de374149b3b12ed2233a9e382f5b658649b7e5afcc70e4aae27189a4356fb080","against":null,"payload":{"description":"Run and inspect the script-owned local two-sibling validation before spending GPU credits.","included_files":[],"name":"verify-environment","skill_md_contents":"---\nname: verify-environment\ndescription: Run and inspect the script-owned local two-sibling validation before spending GPU credits.\n---\n\n# Verify the environment\n\nInspect `main.py`, then run:\n\n```bash\nuv run python main.py validate\n```\n\nThe validation stage must first obtain its example through\n`env.create_dataset(\"train\", Path(\".\"))`, then call\n`castform.validate_environment` once for one real `Environment.run_group`\ncontaining exactly two siblings of that example locally and in the hosted\nsandbox. This exercises the public data materialization and deployment contract\nas part of validation. Hosted validation always runs against the exact assets\nthat were just uploaded — the same ones a launch would train on. Keep the local and hosted rollout-model context\nbudget shared through `VALIDATE_CONFIG[\"max_context_tokens\"]`; the local\nwall-clock backstop is `VALIDATE_CONFIG[\"local_timeout_seconds\"]`.\n\n`validate_environment` first performs static model-parameter checks, then uses\ntracked model sessions locally and remotely to enforce the same sampling and\nmulti-turn history contract as training. Review static and runtime warnings as\nwell as outcomes. A `max_tokens` or `max_completion_tokens` warning is allowed\nwhen the effective cap is acceptable. Sampling conflicts, unsupported controls,\nchanged tools, overlapping generations, and rewritten assistant history are\nerrors. Do not launch while any contract error remains. If validation made only\none model call, treat the “multi-turn history was not exercised” warning as a\ncoverage gap for harnesses expected to loop.\n\n## Read both outcomes\n\nFor each sibling, inspect:\n\n- `termination_reason`;\n- the complete reward mapping and total;\n- evidence that the response was actually scored by the intended reward;\n- any environment, tool, sandbox or judge error logs.\n\nA successful outcome has `termination_reason == \"finished\"` and the reward\ncomponents produced by its scoring hooks. Its scores may legitimately all be\nzero. An operational failure has a different termination reason, no rewards,\nand a visible log entry. It must not cancel the other sibling.\n\nDo not call the baseline green when:\n\n- an outcome failed, even if its reward mapping looks structurally valid;\n- rewards are malformed, non-finite or missing declared keys;\n- the reward is constant for reasons the task does not justify;\n- the judge or verifier failure was mistaken for a valid zero score;\n- group-relative scoring depends on failed siblings or cross-group state.\n\n## Targeted checks\n\nBefore launch, add unit tests for empty, wrong, partial and correct answers. Put\nthem in `tests/` next to `main.py` (its `conftest.py` pins the import path so\n`from main import ...` resolves) and run `uv run pytest tests`. Exercise\ntool exceptions and judge exceptions and assert the failure termination reason,\nzeroed declared shape and log message. For a group-relative reward, verify that one\nfailed sibling does not alter successful siblings' scoring inputs.\n\nIf the environment uses `InjectedAuth(\"judge\")` for the Castform LLM endpoint, Castform validation binds that name to its call-time credential provider for the duration of the run. Rollout\n`model_auth` and named environment bindings are independent; overriding one must\nnot silently override the other. A missing or unknown binding should fail visibly;\nthe environment must not read a platform token itself.\n\nWhen the baseline is green, report both outcomes and ask whether to iterate or load\n**launch-run**. Do not launch automatically.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}