← castformCONTENT HISTORY

Update to castform

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.0.0+codex.20260818184934

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Monitor a Castform run with status, scalar, rollout, and log commands, then diagnose failures or reward drift.",
  "included_files": [],
  "name": "view-progress",
  "skill_md_contents": "---\nname: view-progress\ndescription: Monitor a Castform run with status, scalar, rollout, and log commands, then diagnose failures or reward drift.\n---\n\n# View progress\n\nStart with the run ID printed by `main.py launch`:\n\n```bash\ncastform runs status <run-id>\ncastform runs scalars <run-id> --mode eval --json\ncastform runs logs <run-id>\n```\n\nUse eval, not only train reward, to judge generalization. A rising train curve with\na flat or falling eval curve is overfitting, not success.\n\n## Inspect actual rollouts\n\nScalar totals are not enough to validate a reward. Read transcripts and per-\ncomponent scores in the terminal or as JSON:\n\n```bash\ncastform runs rollouts <run-id> --mode eval\ncastform runs rollout <run-id> <rollout-id>\ncastform runs rollout <run-id> <rollout-id> --json\n```\n\n`runs rollout` can join ground truth from local JSONL. Pass `--dataset <path>`\nwhen the project does not use the default eval/train filenames. Use the text or\nJSON output and the run page printed at launch.\n\nWhen reviewing stored outcomes, distinguish a valid zero score from execution\nfailure using available logs and stored error fields. benchmax produces a non-\n`finished` termination reason locally, but carrying that field faithfully through\nthe trainer and hosted rollout views is pending downstream integration; do not\nclaim a stored run exposes it until the platform does.\n\n## Diagnosis\n\n- `pending` for too long: inspect status and launch/platform logs.\n- `failed` early: inspect environment imports, bundle dependencies and the first\n  rollout error; fix the project and re-run validation before another launch.\n- `stalled`: inspect recent activity and logs; do not infer model quality from an\n  incomplete run.\n- flat rewards: read correct and incorrect transcripts, then test the reward\n  locally for discrimination and accidental bonus paths.\n- eval peaks then declines: record the best step and verify which checkpoint is\n  available before treating the final checkpoint as best.\n- judge errors: fix auth/provider/runtime reliability; never reinterpret the\n  zeroed failure reward as the judge's verdict.\n\n<!-- rag:start -->\nFor RAG, separate retrieval from answer quality: gold never retrieved, gold\nretrieved but not cited, and correct cited answers are different failure modes.\nCheck source-ID canonicalization before changing reward weights.\n<!-- rag:end -->\n\nEvery `runs` read supports `--json`. Use it for programmatic comparison, but do not\nbuild an unbounded polling loop. Re-run status and scalar reads after returning to\na long-running job.\n\nTo cancel a run owned by the current account:\n\n```bash\ncastform stop <run-id>\n```\n\nPreserve the run ID, terminal state and decisive log/reward evidence in the handoff.\n"
}

SHA-256 of public snapshot: 0639ff6dabc34c05a33e44890bcaf0efd9390a9d70f536325295a0fd84655dc0