← castformCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to castform
Snapshot Sep 30, 2026 · 23:14 UTC · version 1.0.0+codex.20260818184934
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Design a benchmax environment, its ordered dataset, tools, and explicit reward shape for a Castform project.",
"included_files": [],
"name": "design-environment",
"skill_md_contents": "---\nname: design-environment\ndescription: Design a benchmax environment, its ordered dataset, tools, and explicit reward shape for a Castform project.\n---\n\n# Design an environment\n\nBefore coding, inspect the maintained [examples](https://github.com/castform-ai/benchmax/tree/main/examples), choose the closest task shape, and follow its `README.md` and `main.py`.\n\nUse `BaseEnv` for the standard OpenAI-compatible chat and tool loop. Use\n`HarborEnv` for a Harbor dataset/package or tasks that ship their own instruction,\nsandbox, and verifier. Start with `aime` for packaged tasks or `harvey` for a\ncustom harness. Do not infer Harbor from a judge, tool, or Dockerfile alone.\nExtend `Environment` directly only for another genuinely different rollout loop.\n\n## Required BaseEnv shape\n\n```python\nfrom pathlib import Path\n\nfrom benchmax.envs import BaseEnv, BaseRollout, DatasetSplit, JsonlDataset\nfrom benchmax.envs.base import resolve_dataset_path\nfrom benchmax.rewards import extract_completion_text\n\n\nclass MyEnv(BaseEnv):\n max_turns = 1\n\n async def create_dataset(\n self, split: DatasetSplit, base_dir: Path\n ) -> JsonlDataset:\n path = resolve_dataset_path(base_dir, f\"{split}.jsonl\")\n return JsonlDataset(path, row_to_example=...)\n\n async def compute_reward(self, rollout: BaseRollout) -> dict[str, float]:\n answer = extract_completion_text(rollout.messages)\n return {\"correct\": float(answer == rollout.example_args[\"answer\"])}\n```\n\nBuild each `Example` with a stable ID, normally\n`canonical_example_id(payload)`. `prompt_messages` is the reserved BaseEnv payload\nfield; all other fields become `rollout.example_args`. Put system messages in\n`prompt_messages` rather than ambient module state.\n\n`Dataset` is an ordered base class, not a cleaning pipeline. The environment owns\nthe runtime representation and can store lightweight references in payloads.\nPreparation, cleaning and QA generation belong in the project data script.\n\n## Reward contract\n\n- Return named reward components from successful individual and group reward hooks.\n- Return finite numbers and make correctness the dominant signal.\n- Let judge, model, tool and sandbox operational failures propagate through the\n typed runtime path. benchmax logs them and returns no rewards with a\n non-`finished` termination reason.\n- Do not catch a judge failure and report it as a legitimate score.\n- Programming, malformed-result and configuration errors should remain loud.\n\nOverride `compute_group_rewards` only when scoring genuinely depends on successful\nsiblings. Failed siblings are excluded from group-relative scoring, and a failed\ngroup judge zeroes otherwise-successful siblings without cancelling the group.\n\n## Optional tools\n\n`BaseEnv` supplies no tools by default. A tool-using environment returns standard\nOpenAI tool schemas from `list_tools` and dispatches them in `run_tool`:\n\n```python\nasync def run_tool(self, rollout_id: str, tool_name: str, **tool_args):\n if tool_name != \"lookup\":\n raise ValueError(f\"unknown tool: {tool_name}\")\n return await self.lookup(tool_args[\"query\"])\n```\n\nKeep clients pickle-safe. Use `InjectedAuth` for calls through the Castform LLM endpoint so Castform supplies the current session credential. Use explicit `StaticBearerAuth` for a user-managed external endpoint; never read Castform credentials from benchmax environment code.\n\n## Model-request ownership\n\nTreat model sampling as trainer-owned. A harness may request an output ceiling\nwith `max_tokens` or `max_completion_tokens`; static validation emits a warning\nbecause Castform may clamp that ceiling to the remaining context budget. Do not\nset `temperature`, `top_p`, `top_k`, penalties, `seed`, or `stop` in agent or\nnested model kwargs. Static validation rejects them instead of allowing a later\ntraining failure. It also rejects unsupported controls such as `n > 1`, forced\n`tool_choice`, logprobs, and non-text response formats.\n\n## Review before handoff\n\n1. Test dataset identity and split ordering.\n2. Unit-test empty, wrong, partial and correct completions against every reward key.\n3. Exercise tool errors and judge errors and confirm zero rewards plus an explicit\n termination reason and log.\n4. Load **verify-environment** and run the real two-sibling validation.\n5. Record every remote runtime import in `RUNTIME_DEPENDENCIES` for **launch-run**.\n\nFor Harbor, require sandbox credentials and the matching provider extra in the\nbundle dependencies, for example\n`harbor[modal]>=0.18,<0.19`.\n"
}SHA-256 of public snapshot: b6be970d445965fda119ad6c953f72334c7579a281a7d9090da0b8158179d079