← Plugin catalog
Developer Tools
castform
castform v1.0.0+codex.20260818184934
Publisher description
From the marketplace listing
design, validate, launch, and monitor castform post-training runs from codex.
Language: English · Automatically detected from descriptions.
Publisher keywords
Search terms declared by the publisher.
Files & skills
File archives
Plugin package12 files · 160 KBBrowse files →
Skill instructions
design-environment4.42 KB
---
name: design-environment
description: Design a benchmax environment, its ordered dataset, tools, and explicit reward shape for a Castform project.
---
# Design an environment
Before coding, inspect the maintained [examples](https://github.com/castform-ai/benchmax/tree/main/examples), choose the closest task shape, and follow its `README.md` and `main.py`.
Use `BaseEnv` for the standard OpenAI-compatible chat and tool loop. Use
`HarborEnv` for a Harbor dataset/package or tasks that ship their own instruction,
sandbox, and verifier. Start with `aime` for packaged tasks or `harvey` for a
custom harness. Do not infer Harbor from a judge, tool, or Dockerfile alone.
Extend `Environment` directly only for another genuinely different rollout loop.
## Required BaseEnv shape
```python
from pathlib import Path
from benchmax.envs import BaseEnv, BaseRollout, DatasetSplit, JsonlDataset
from benchmax.envs.base import resolve_dataset_path
from benchmax.rewards import extract_completion_text
class MyEnv(BaseEnv):
max_turns = 1
async def create_dataset(
self, split: DatasetSplit, base_dir: Path
) -> JsonlDataset:
path = resolve_dataset_path(base_dir, f"{split}.jsonl")
return JsonlDataset(path, row_to_example=...)
async def compute_reward(self, rollout: BaseRollout) -> dict[str, float]:
answer = extract_completion_text(rollout.messages)
return {"correct": float(answer == rollout.example_args["answer"])}
```
Build each `Example` with a stable ID, normally
`canonical_example_id(payload)`. `prompt_messages` is the reserved BaseEnv payload
field; all other fields become `rollout.example_args`. Put system messages in
`prompt_messages` rather than ambient module state.
`Dataset` is an ordered base class, not a cleaning pipeline. The environment owns
the runtime representation and can store lightweight references in payloads.
Preparation, cleaning and QA generation belong in the project data script.
## Reward contract
- Return named reward components from successful individual and group reward hooks.
- Return finite numbers and make correctness the dominant signal.
- Let judge, model, tool and sandbox operational failures propagate through the
typed runtime path. benchmax logs them and returns no rewards with a
non-`finished` termination reason.
- Do not catch a judge failure and report it as a legitimate score.
- Programming, malformed-result and configuration errors should remain loud.
Override `compute_group_rewards` only when scoring genuinely depends on successful
siblings. Failed siblings are excluded from group-relative scoring, and a failed
group judge zeroes otherwise-successful siblings without cancelling the group.
## Optional tools
`BaseEnv` supplies no tools by default. A tool-using environment returns standard
OpenAI tool schemas from `list_tools` and dispatches them in `run_tool`:
```python
async def run_tool(self, rollout_id: str, tool_name: str, **tool_args):
if tool_name != "lookup":
raise ValueError(f"unknown tool: {tool_name}")
return await self.lookup(tool_args["query"])
```
Keep clients pickle-safe. Use `InjectedAuth` for calls through the Castform LLM endpoint so Castform supplies the current session credential. Use explicit `StaticBearerAuth` for a user-managed external endpoint; never read Castform credentials from benchmax environment code.
## Model-request ownership
Treat model sampling as trainer-owned. A harness may request an output ceiling
with `max_tokens` or `max_completion_tokens`; static validation emits a warning
because Castform may clamp that ceiling to the remaining context budget. Do not
set `temperature`, `top_p`, `top_k`, penalties, `seed`, or `stop` in agent or
nested model kwargs. Static validation rejects them instead of allowing a later
training failure. It also rejects unsupported controls such as `n > 1`, forced
`tool_choice`, logprobs, and non-text response formats.
## Review before handoff
1. Test dataset identity and split ordering.
2. Unit-test empty, wrong, partial and correct completions against every reward key.
3. Exercise tool errors and judge errors and confirm zero rewards plus an explicit
termination reason and log.
4. Load **verify-environment** and run the real two-sibling validation.
5. Record every remote runtime import in `RUNTIME_DEPENDENCIES` for **launch-run**.
For Harbor, require sandbox credentials and the matching provider extra in the
bundle dependencies, for example
`harbor[modal]>=0.18,<0.19`.
generate-data3.94 KB
---
name: generate-data
description: Prepare, clean, generate, or reference the data used by a Castform environment from project-owned Python scripts.
---
# Generate or reference data
Data preparation is a Python workflow. Keep it in `main.py`'s `generate_data`
stage or a nearby project script so its inputs, transformations and outputs are
reviewable. There is no separate data or corpus orchestration CLI.
## BaseEnv JSONL path
The generated BaseEnv seed reads `train.jsonl` and `eval.jsonl`. Each line is a
JSON object containing the fields the environment's row converter and reward
actually use:
```jsonl
{"prompt": "What is 2 + 2?", "ground_truth": "4"}
```
- Keep train and eval disjoint.
- Build stable example identity from semantic payload content, not row position or
machine-local paths.
- Start small enough to inspect manually, then expand after validation shows a
meaningful signal.
- Keep large integer identifiers as strings when data will cross JSON/JavaScript
boundaries.
- Record provenance and make regeneration idempotent; never overwrite curated data
without an explicit force flag.
`upload_assets` accepts optional train and eval rows. The launch script
decides what it uploads: omit a split when the environment resolves it at runtime,
and do not upload unrelated preparation artifacts. `None` means “do not upload”;
an empty list deliberately uploads an empty JSONL.
<!-- rag:start -->
## Hosted corpus and RAG
Install `castform[rag]` in the project and use public modules under
`castform.rag` from the data stage. Before implementing the workflow, inspect the
matching maintained example:
- `neon_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/neon_rag
- `turbopuffer_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/turbopuffer_rag
- `chroma_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/chroma_rag
- `pinecone_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/pinecone_rag
Use its `README.md`, `main.py`, `data.py`, `environment.py`, and `search.py` as the
reference for the provider. Typical data code composes:
- `castform.rag.chunkers` to turn source files into chunks;
- `castform.rag.corpus.postgres.client.CorpusClient` to create/find a corpus and
upload chunks;
- `castform.rag.qa_generation` to build grounded QA rows;
- the provider's example-local search adapter for runtime reads.
Read the concrete class signatures before wiring them; these are library
components, not one magical pipeline command. Persist generated rows as ordinary
project data and test at least one known retrieval query before validation.
Replace the generic seed environment and rows with the selected RAG example's
structure. Confirm that each row contains `question`, `answer`, and
`reference_chunks`, and that every reference chunk carries the source metadata
expected by the citation reward.
<!-- rag:end -->
## Harbor-managed datasets
A Harbor dataset may be a local directory, Harbor package, registry reference or
Git repository resolved by `HarborEnv` during runtime. Do not duplicate that data
into JSONL merely to match the BaseEnv example. Only add an upload step when the
chosen Harbor workflow genuinely needs a folder or artifact uploaded.
## Traces
Normalize provider traces with the adapter for that provider and pass them through
`castform.traces.TracesPipeline`, which ships with the base `castform` package.
Keep the resulting train/eval split and detected prompt/tool assumptions visible
in the project.
Inspect for secrets, relayed tool output, duplicates and trivial examples before
using traces as training data.
## Verification
Before handing off to **verify-environment**:
1. validate the output schema with the environment's row converter;
2. check stable IDs and train/eval overlap;
3. inspect representative easy, hard and malformed rows;
4. confirm any corpus or Git reference is reachable in the rollout runtime;
5. run the deterministic reward against known answers where possible.
launch-run3.88 KB
---
name: launch-run
description: Review, bundle, upload, and explicitly launch a Castform GPU training run from the project script.
---
# Launch a run
Use this only after **verify-environment** reports a believable green baseline.
Launching spends GPU credits. The workflow lives in `main.py`, not a CLI launch
command:
```bash
uv run python main.py launch
```
Do not pass `--yes` unless the user has already explicitly authorized the cost.
Never launch after `validate_environment` reports a static or runtime sampling
or history-contract error. Review output-cap warnings (`max_tokens` or
`max_completion_tokens`) and confirm any effective clamp is intentional.
## Required ordering
Accept user configuration as explicit `main.py` arguments, normalize it once in
`_constructor_args(args)`, and reuse that dictionary for local construction and
`dump_bundle`. Avoid ambient `os.environ` reads in environments, tools, rewards,
and harness configuration.
Read the script and confirm that the launch action does all of the following
in order:
1. builds one `Bundle` with `dump_bundle`;
2. passes that exact object to `upload_assets(bundle=bundle, ...)`;
3. validates the uploaded assets (locally and in the hosted sandbox) and stops
on failure;
4. asks the human to confirm a credit-spending GPU launch;
5. passes the same uploaded paths to `TrainerClient.launch_training_run` — the
run trains on precisely what was validated.
The upload helper must not silently rebundle the environment, and launch must
not re-upload.
Dataset upload is explicit and optional. Supply `train_dataset` and/or
`eval_dataset` only for splits Castform should upload. Omit them for data resolved
by the environment at runtime (for example Harbor- or Git-managed data). Do not
use an empty list as an omission sentinel: it uploads an empty JSONL file.
## Dependencies
`RUNTIME_DEPENDENCIES` is explicit and authoritative for the remote rollout
runtime:
```python
bundle = dump_bundle(
CustomEnv,
constructor_args=constructor_args,
pip_dependencies=RUNTIME_DEPENDENCIES,
)
```
List every external package imported while the environment, tools or rewards run.
Do not copy the whole project dependency list automatically: data-preparation and
development packages may not belong in the rollout image. benchmax captures local
modules under the environment project automatically. Source from another project
must be explicit: use `local_modules=` to capture it, or list its installed
distribution in `pip_dependencies` to keep it as a remote reference.
For Harbor, add the selected provider extra explicitly, such as
`harbor[modal]>=0.18,<0.19` or `harbor[daytona]>=0.18,<0.19`.
## Launch configuration
Review `LAUNCH_CONFIG` in source. In particular:
- `max_context_tokens` is the whole-rollout prompt-plus-response token budget;
- keep trainer turn/tool limits compatible with the environment's own limits;
- start with modest epochs and judge the eval curve, not only train reward;
- use `TrainerClient.list_launch_args()` when you need the live accepted schema
instead of guessing an argument name.
<!-- rag:start -->
For search environments, budget for repeated tool output across turns. Confirm
the rollout bundle includes the runtime search client but not large local corpus-
preparation dependencies unless the environment imports them.
<!-- rag:end -->
## Credentials
Use `InjectedAuth` for model and judge calls through the Castform LLM endpoint so the hosted runtime supplies the current Castform credential. User-managed external endpoints use explicit `StaticBearerAuth`. Harbor sandbox credentials are currently explicit constructor inputs. Review static credentials before bundling and limit their scope.
## Handoff
Record the run ID printed by the script, then load **view-progress**. If upload or
launch fails, preserve the error, correct the script or credentials, and rerun the
smallest failed stage. Never bypass a failed validation gate.
setup1023 Bytes
--- name: setup description: One-time bootstrap for a new Castform project — install the Castform CLI and scaffold the Codex guide, skills, and starter environment. Run this before any other Castform skill. --- # Set up Castform The other Castform skills (`design-environment`, `generate-data`, `verify-environment`, `launch-run`, `view-progress`, and `train-sft`) assume the CLI is installed and the current project has been scaffolded. 1. Install or upgrade the CLI: ```bash uv tool install -U castform ``` 2. Scaffold the current project for Codex: ```bash castform setup --agent codex ``` This signs in through the browser when necessary, then writes `AGENTS.md`, `GETTING_STARTED.md`, the per-stage skills, and a runnable starter environment. It is safe to rerun because existing files are skipped unless `--force` is passed. 3. Confirm `castform --version` succeeds and the project contains `AGENTS.md` plus `.agents/skills/`. Continue with the `design-environment` skill.
train-sft6.32 KB
---
name: train-sft
description: Build a validated benchmax-sft-v1 dataset from existing labeled conversations, upload it, and explicitly launch a supervised finetuning run — no environment, rewards, or rollouts.
---
# Train with SFT
Use this when the user **already has the completions they want the model to
imitate** — support transcripts, labeled chat data, tool-call traces, input→output
pairs. There is no environment, tool loop, reward, or validate stage: the RL loop
in this project's other skills does not apply. If the task needs the model to
*discover* good behavior against a scorer, use the RL skills instead.
## The whole flow
```python
from benchmax.sft import SftDataset, SftDatasetError
from castform.platform import SftTrainingConfig, TrainerClient, upload_sft_assets
train = SftDataset.from_jsonl("train.jsonl") # or SftDataset.from_rows(rows)
uploaded = upload_sft_assets(dataset=train, run_name="support-sft")
run_id = TrainerClient().launch_sft_run(
assets=uploaded,
name="support-sft",
config=SftTrainingConfig(num_epochs=1, learning_rate=1e-5, seed=42),
)
```
Construction is **all-or-nothing**: `SftDataset` either satisfies the whole
`benchmax-sft-v1` contract or raises `SftDatasetError` with every issue, ordered
and line-aware. Fix the data the diagnostics point at — never pre-filter rows
silently or patch around individual issues without telling the user.
## Row contract (one JSON object per line)
```json
{
"messages": [
{"role": "user", "content": "What is 2 + 2?"},
{"role": "assistant", "content": "4", "weight": 1}
],
"tools": [],
"metadata": {"id": "optional producer identity"}
}
```
- `system`/`user`: exactly `role` + non-empty string `content`.
- `assistant`: optional string-or-null `content`, optional non-empty
`tool_calls`, optional integer `weight` `0 | 1` (omitted means `1`). Each
assistant turn needs content or a tool call; each row needs at least one
assistant turn with effective weight `1`.
- `tool` results: exactly `role` + string `content` + non-empty `tool_call_id`;
every tool call gets exactly one result, in declaration order, before the
next non-tool message. Tool definitions are OpenAI function shapes;
`function.arguments` must decode as a JSON object.
- Everything else is rejected: images/audio/multimodal parts, fractional
weights, legacy prompt/completion keys, unknown fields, duplicate JSON keys,
rows over 1 MiB, more than 1024 messages.
**Masking:** set `weight: 0` on assistant turns that are context, not target
(earlier drafts, retrieved answers, another model's output). Only weight-1
turns contribute to the loss.
## What the user actually chooses
| arg | accepted | default |
|---|---|---|
| `num_epochs` | 1–100 | 1 |
| `learning_rate` | (0, 0.1] | 1e-5 |
| `max_context_tokens` | 256–8192, or 32768 / 65536 / 131072 (long context — must be enabled on the platform) | 8192 |
| `save_interval` (steps) | 1–10000 | 20 |
| `seed` | 0–2147483647 | 42 |
| `lr_decay_style` | `"constant"` or `"cosine"` | unset — keeps the platform default |
| `min_lr` | ≥ 0 and below `learning_rate` | unset |
| `warmup_ratio` | 0–0.5 of total steps | unset |
| `adam_beta2` | 0.9–0.999 | unset |
| `grad_clip` | (0, 10] | unset |
| `lora_rank` | 32 or 64 | unset — keeps the platform's fixed policy |
| `global_batch_size` | 4–64, multiple of 4 at `max_context_tokens` ≤ 8192; 1–64 at a long-context rung (the platform picks the topology, so it decides divisibility) | unset — 4 |
| `eval_interval` (steps) | 1–10000; only with an eval set | unset — derived from `save_interval` |
Every arg whose default reads `unset` is optional: leave it out and the
platform's own value applies, so an untouched config behaves exactly as before
these knobs existed. Do not set one just to restate a default.
`lora_rank` picks the adapter's rank; alpha is always derived as 2x and is not
a knob. Rank 128 is not accepted — it trains, but serving cannot load it.
Model (`Qwen/Qwen3.5-4B`) and GPU topology are platform-owned — do not invent
knobs for them. A row that renders past `max_context_tokens` tokens fails the
run's preflight; trim long rows up front.
## Held-out eval (optional)
Pass a second dataset to score during training:
```python
assets = upload_sft_assets(dataset=train, eval_dataset=held_out, run_name="...")
```
The eval set is scored on the live weights between training steps and plotted
as `eval/loss` beside `train/loss`. Rules worth knowing before you offer it:
- **At most 2048 rows**, under the same per-row limits as training rows. The
bound is a compute bound, not a taste one — eval runs on the training GPUs
and pauses training while it does.
- **It is never billed.** Eval forward passes cost the user nothing, which is
also why the size and cadence are capped rather than left open.
- **Cadence defaults to `save_interval`**, so by default every eval lands on a
checkpoint. Set `eval_interval` only to make it sparser; a much denser
cadence is rejected at launch.
- An eval set is part of the data identity: changing it produces a different
upload prefix, and a resumed run must keep the one it started with.
Skip it for small or exploratory runs — a held-out split costs training rows,
and `train/loss` alone answers "is this learning at all".
## Cost and consent
`launch_sft_run` spends GPU credits. Ask the human before calling it, every
time — preparing and uploading the dataset first is free and fine. Steps per
epoch ≈ rows / 4 (tiny datasets pad by repeating their first rows, so 1–3-row
datasets train on repeats; prefer at least a few dozen rows). Stopping a run
keeps only checkpoints already uploaded; work since the last one is lost.
If launch fails with "SFT launch is not enabled", the platform gate is off for
this account — surface that to the user rather than retrying.
## Monitor
Same as any run (`view-progress` skill): `castform runs status <id>`, and
`castform runs scalars <id>` — watch `train/loss` fall, and `eval/loss` too
when the run has an eval set. There are no rollouts or reward curves for SFT
runs. The run page shows loss, dataset prefix, and config; runs with an eval
set also get an eval tab, chart-only.
## Reference
The canonical worked example (streaming a pinned public corpus, bounded
mapping, offline tests, explicit paid `--launch` gate) ships in the benchmax
repo under `examples/sft/pii_masking/`.
verify-environment3.45 KB
---
name: verify-environment
description: Run and inspect the script-owned local two-sibling validation before spending GPU credits.
---
# Verify the environment
Inspect `main.py`, then run:
```bash
uv run python main.py validate
```
The validation stage must first obtain its example through
`env.create_dataset("train", Path("."))`, then call
`castform.validate_environment` once for one real `Environment.run_group`
containing exactly two siblings of that example locally and in the hosted
sandbox. This exercises the public data materialization and deployment contract
as part of validation. Hosted validation always runs against the exact assets
that were just uploaded — the same ones a launch would train on. Keep the local and hosted rollout-model context
budget shared through `VALIDATE_CONFIG["max_context_tokens"]`; the local
wall-clock backstop is `VALIDATE_CONFIG["local_timeout_seconds"]`.
`validate_environment` first performs static model-parameter checks, then uses
tracked model sessions locally and remotely to enforce the same sampling and
multi-turn history contract as training. Review static and runtime warnings as
well as outcomes. A `max_tokens` or `max_completion_tokens` warning is allowed
when the effective cap is acceptable. Sampling conflicts, unsupported controls,
changed tools, overlapping generations, and rewritten assistant history are
errors. Do not launch while any contract error remains. If validation made only
one model call, treat the “multi-turn history was not exercised” warning as a
coverage gap for harnesses expected to loop.
## Read both outcomes
For each sibling, inspect:
- `termination_reason`;
- the complete reward mapping and total;
- evidence that the response was actually scored by the intended reward;
- any environment, tool, sandbox or judge error logs.
A successful outcome has `termination_reason == "finished"` and the reward
components produced by its scoring hooks. Its scores may legitimately all be
zero. An operational failure has a different termination reason, no rewards,
and a visible log entry. It must not cancel the other sibling.
Do not call the baseline green when:
- an outcome failed, even if its reward mapping looks structurally valid;
- rewards are malformed, non-finite or missing declared keys;
- the reward is constant for reasons the task does not justify;
- the judge or verifier failure was mistaken for a valid zero score;
- group-relative scoring depends on failed siblings or cross-group state.
## Targeted checks
Before launch, add unit tests for empty, wrong, partial and correct answers. Put
them in `tests/` next to `main.py` (its `conftest.py` pins the import path so
`from main import ...` resolves) and run `uv run pytest tests`. Exercise
tool exceptions and judge exceptions and assert the failure termination reason,
zeroed declared shape and log message. For a group-relative reward, verify that one
failed sibling does not alter successful siblings' scoring inputs.
If the environment uses `InjectedAuth("judge")` for the Castform LLM endpoint, Castform validation binds that name to its call-time credential provider for the duration of the run. Rollout
`model_auth` and named environment bindings are independent; overriding one must
not silently override the other. A missing or unknown binding should fail visibly;
the environment must not read a platform token itself.
When the baseline is green, report both outcomes and ask whether to iterate or load
**launch-run**. Do not launch automatically.
view-progress2.66 KB
--- name: view-progress description: Monitor a Castform run with status, scalar, rollout, and log commands, then diagnose failures or reward drift. --- # View progress Start with the run ID printed by `main.py launch`: ```bash castform runs status <run-id> castform runs scalars <run-id> --mode eval --json castform runs logs <run-id> ``` Use eval, not only train reward, to judge generalization. A rising train curve with a flat or falling eval curve is overfitting, not success. ## Inspect actual rollouts Scalar totals are not enough to validate a reward. Read transcripts and per- component scores in the terminal or as JSON: ```bash castform runs rollouts <run-id> --mode eval castform runs rollout <run-id> <rollout-id> castform runs rollout <run-id> <rollout-id> --json ``` `runs rollout` can join ground truth from local JSONL. Pass `--dataset <path>` when the project does not use the default eval/train filenames. Use the text or JSON output and the run page printed at launch. When reviewing stored outcomes, distinguish a valid zero score from execution failure using available logs and stored error fields. benchmax produces a non- `finished` termination reason locally, but carrying that field faithfully through the trainer and hosted rollout views is pending downstream integration; do not claim a stored run exposes it until the platform does. ## Diagnosis - `pending` for too long: inspect status and launch/platform logs. - `failed` early: inspect environment imports, bundle dependencies and the first rollout error; fix the project and re-run validation before another launch. - `stalled`: inspect recent activity and logs; do not infer model quality from an incomplete run. - flat rewards: read correct and incorrect transcripts, then test the reward locally for discrimination and accidental bonus paths. - eval peaks then declines: record the best step and verify which checkpoint is available before treating the final checkpoint as best. - judge errors: fix auth/provider/runtime reliability; never reinterpret the zeroed failure reward as the judge's verdict. <!-- rag:start --> For RAG, separate retrieval from answer quality: gold never retrieved, gold retrieved but not cited, and correct cited answers are different failure modes. Check source-ID canonicalization before changing reward weights. <!-- rag:end --> Every `runs` read supports `--json`. Use it for programmatic comparison, but do not build an unbounded polling loop. Re-run status and scalar reads after returning to a long-running job. To cancel a run owned by the current account: ```bash castform stop <run-id> ``` Preserve the run ID, terminal state and decisive log/reward evidence in the handoff.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- castform
- Keywords
- See publisher keywords
Declared capabilities
- Interactive
- Write
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 3, 2026 · 00:00 UTC
- Collection status
- Collected
plugins_6a84a979b5a081918922e037047514b5
Download plugin data (JSON)