← castformCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to castform
Snapshot Sep 30, 2026 · 23:14 UTC · version 1.0.0+codex.20260818184934
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Monitor a Castform run with status, scalar, rollout, and log commands, then diagnose failures or reward drift.",
"included_files": [],
"name": "view-progress",
"skill_md_contents": "---\nname: view-progress\ndescription: Monitor a Castform run with status, scalar, rollout, and log commands, then diagnose failures or reward drift.\n---\n\n# View progress\n\nStart with the run ID printed by `main.py launch`:\n\n```bash\ncastform runs status <run-id>\ncastform runs scalars <run-id> --mode eval --json\ncastform runs logs <run-id>\n```\n\nUse eval, not only train reward, to judge generalization. A rising train curve with\na flat or falling eval curve is overfitting, not success.\n\n## Inspect actual rollouts\n\nScalar totals are not enough to validate a reward. Read transcripts and per-\ncomponent scores in the terminal or as JSON:\n\n```bash\ncastform runs rollouts <run-id> --mode eval\ncastform runs rollout <run-id> <rollout-id>\ncastform runs rollout <run-id> <rollout-id> --json\n```\n\n`runs rollout` can join ground truth from local JSONL. Pass `--dataset <path>`\nwhen the project does not use the default eval/train filenames. Use the text or\nJSON output and the run page printed at launch.\n\nWhen reviewing stored outcomes, distinguish a valid zero score from execution\nfailure using available logs and stored error fields. benchmax produces a non-\n`finished` termination reason locally, but carrying that field faithfully through\nthe trainer and hosted rollout views is pending downstream integration; do not\nclaim a stored run exposes it until the platform does.\n\n## Diagnosis\n\n- `pending` for too long: inspect status and launch/platform logs.\n- `failed` early: inspect environment imports, bundle dependencies and the first\n rollout error; fix the project and re-run validation before another launch.\n- `stalled`: inspect recent activity and logs; do not infer model quality from an\n incomplete run.\n- flat rewards: read correct and incorrect transcripts, then test the reward\n locally for discrimination and accidental bonus paths.\n- eval peaks then declines: record the best step and verify which checkpoint is\n available before treating the final checkpoint as best.\n- judge errors: fix auth/provider/runtime reliability; never reinterpret the\n zeroed failure reward as the judge's verdict.\n\n<!-- rag:start -->\nFor RAG, separate retrieval from answer quality: gold never retrieved, gold\nretrieved but not cited, and correct cited answers are different failure modes.\nCheck source-ID canonicalization before changing reward weights.\n<!-- rag:end -->\n\nEvery `runs` read supports `--json`. Use it for programmatic comparison, but do not\nbuild an unbounded polling loop. Re-run status and scalar reads after returning to\na long-running job.\n\nTo cancel a run owned by the current account:\n\n```bash\ncastform stop <run-id>\n```\n\nPreserve the run ID, terminal state and decisive log/reward evidence in the handoff.\n"
}SHA-256 of public snapshot: 0639ff6dabc34c05a33e44890bcaf0efd9390a9d70f536325295a0fd84655dc0