{"id":18073,"plugin_id":"plugins_6a84a979b5a081918922e037047514b5","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:36.262Z","digest":"4133b5823312be8d8b8e0f9ca4f5e39d188cead999ea0d421a580f780a50d764","against":null,"payload":{"description":"Prepare, clean, generate, or reference the data used by a Castform environment from project-owned Python scripts.","included_files":[],"name":"generate-data","skill_md_contents":"---\nname: generate-data\ndescription: Prepare, clean, generate, or reference the data used by a Castform environment from project-owned Python scripts.\n---\n\n# Generate or reference data\n\nData preparation is a Python workflow. Keep it in `main.py`'s `generate_data`\nstage or a nearby project script so its inputs, transformations and outputs are\nreviewable. There is no separate data or corpus orchestration CLI.\n\n## BaseEnv JSONL path\n\nThe generated BaseEnv seed reads `train.jsonl` and `eval.jsonl`. Each line is a\nJSON object containing the fields the environment's row converter and reward\nactually use:\n\n```jsonl\n{\"prompt\": \"What is 2 + 2?\", \"ground_truth\": \"4\"}\n```\n\n- Keep train and eval disjoint.\n- Build stable example identity from semantic payload content, not row position or\n  machine-local paths.\n- Start small enough to inspect manually, then expand after validation shows a\n  meaningful signal.\n- Keep large integer identifiers as strings when data will cross JSON/JavaScript\n  boundaries.\n- Record provenance and make regeneration idempotent; never overwrite curated data\n  without an explicit force flag.\n\n`upload_assets` accepts optional train and eval rows. The launch script\ndecides what it uploads: omit a split when the environment resolves it at runtime,\nand do not upload unrelated preparation artifacts. `None` means “do not upload”;\nan empty list deliberately uploads an empty JSONL.\n\n<!-- rag:start -->\n## Hosted corpus and RAG\n\nInstall `castform[rag]` in the project and use public modules under\n`castform.rag` from the data stage. Before implementing the workflow, inspect the\nmatching maintained example:\n\n- `neon_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/neon_rag\n- `turbopuffer_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/turbopuffer_rag\n- `chroma_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/chroma_rag\n- `pinecone_rag`: https://github.com/castform-ai/benchmax/tree/main/examples/pinecone_rag\n\nUse its `README.md`, `main.py`, `data.py`, `environment.py`, and `search.py` as the\nreference for the provider. Typical data code composes:\n\n- `castform.rag.chunkers` to turn source files into chunks;\n- `castform.rag.corpus.postgres.client.CorpusClient` to create/find a corpus and\n  upload chunks;\n- `castform.rag.qa_generation` to build grounded QA rows;\n- the provider's example-local search adapter for runtime reads.\n\nRead the concrete class signatures before wiring them; these are library\ncomponents, not one magical pipeline command. Persist generated rows as ordinary\nproject data and test at least one known retrieval query before validation.\n\nReplace the generic seed environment and rows with the selected RAG example's\nstructure. Confirm that each row contains `question`, `answer`, and\n`reference_chunks`, and that every reference chunk carries the source metadata\nexpected by the citation reward.\n<!-- rag:end -->\n\n## Harbor-managed datasets\n\nA Harbor dataset may be a local directory, Harbor package, registry reference or\nGit repository resolved by `HarborEnv` during runtime. Do not duplicate that data\ninto JSONL merely to match the BaseEnv example. Only add an upload step when the\nchosen Harbor workflow genuinely needs a folder or artifact uploaded.\n\n## Traces\n\nNormalize provider traces with the adapter for that provider and pass them through\n`castform.traces.TracesPipeline`, which ships with the base `castform` package.\nKeep the resulting train/eval split and detected prompt/tool assumptions visible\nin the project.\nInspect for secrets, relayed tool output, duplicates and trivial examples before\nusing traces as training data.\n\n## Verification\n\nBefore handing off to **verify-environment**:\n\n1. validate the output schema with the environment's row converter;\n2. check stable IDs and train/eval overlap;\n3. inspect representative easy, hard and malformed rows;\n4. confirm any corpus or Git reference is reachable in the rollout runtime;\n5. run the deterministic reward against known answers where possible.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}