{"id":22257,"plugin_id":"plugins_6ab026cc7f308191ba091a7edaa9d5ed","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:17:12.485Z","digest":"3caaa3550d4e08bb61ce5ebfdf3b7ccbf870da79fccab3f6e2c3b091d3ac034d","against":null,"payload":{"description":"Use when the user asks to design, evaluate, troubleshoot, or operate a continual-learning or self-improvement workflow for AI agents using Reef.","included_files":[{"relative_path":"SOURCE-METADATA.yaml","size_in_bytes":58},{"relative_path":"agents/openai.yaml","size_in_bytes":247},{"relative_path":"references/LICENSE","size_in_bytes":11337},{"relative_path":"references/README.md","size_in_bytes":17578},{"relative_path":"references/cli.rst","size_in_bytes":7458},{"relative_path":"references/configuration.rst","size_in_bytes":50970},{"relative_path":"references/core-loop.rst","size_in_bytes":2633},{"relative_path":"references/harness-adapters.rst","size_in_bytes":44226},{"relative_path":"references/installation.rst","size_in_bytes":2593},{"relative_path":"references/intro.rst","size_in_bytes":5159},{"relative_path":"references/quickstart.rst","size_in_bytes":8149},{"relative_path":"references/write-a-harness-method.rst","size_in_bytes":10875},{"relative_path":"references/write-a-recipe.rst","size_in_bytes":11704}],"name":"reef-agent-improvement","skill_md_contents":"---\nname: reef-agent-improvement\ndescription: Use when the user asks to design, evaluate, troubleshoot, or operate a continual-learning or self-improvement\n  workflow for AI agents using Reef.\nlicense: Apache-2.0\ncompatibility: Guidance works anywhere; executing Reef requires Python 3.10+, the reef-infra runtime, and any additional dependencies\n  required by the chosen recipe.\n---\n\n# Reef Agent Improvement\n\n## Overview\n\nUse this skill to work with **Reef**, the continual-learning infrastructure for self-improving agents. Reef connects live inference, recorded interactions, feedback, learning/update jobs, evaluation, and versioned artifact delivery.\n\nTreat the bundled reference files as the source of truth for Reef-specific commands and configuration. Do not infer undocumented flags, APIs, or installation state.\n\n## When to Use\n\nUse Reef-oriented guidance when the task involves one or more of these needs:\n\n- improving an agent harness over repeated interactions;\n- learning from scored or structured feedback;\n- evolving prompts, rules, skills, or other harness artifacts;\n- model-weight training through a supported recipe;\n- test-time training or search with a measurable objective;\n- managing versioned agent artifacts and accepted updates;\n- integrating a harness adapter, processor, executor, recipe, or evaluation policy;\n- diagnosing a Reef deployment, configuration, or feedback pipeline.\n\nDo not introduce Reef for a one-off prompt edit, a simple static agent, or a task with no repeated feedback/evaluation loop unless the user explicitly asks for Reef.\n\n## Runtime Check\n\nBefore giving commands that assume Reef is installed, inspect the environment when tool access is available. Suitable checks include:\n\n```bash\npython --version\npython -c \"import reef; print(getattr(reef, '__version__', 'reef import OK'))\"\nreef --help\n```\n\nIf the runtime is unavailable, explain that the plugin supplies **guidance and references**, not the Reef package itself. Use `references/installation.rst` and `references/quickstart.rst` for setup guidance rather than pretending execution succeeded.\n\n## Workflow Selection\n\nDetermine which Reef surface matches the user's goal before proposing configuration:\n\n1. **Harness optimization** — prompts, rules, skills, adapters, or other agent artifacts improve from representative tasks plus evaluation. This generally does not require local training GPUs.\n2. **Model-weight training** — a supported trainable model and training stack update weights from eligible feedback.\n3. **Test-time training / scientific search** — an execution environment and correctness or objective function drive iterative improvement.\n\nWhen uncertain, ask what is being improved, how success is measured, and what feedback is available.\n\n## Core Loop\n\nReason about Reef systems using its four-stage loop:\n\n- **Serve:** handle requests and record interactions.\n- **Observe:** associate feedback with recorded interactions and determine eligibility.\n- **Grow:** generate candidate updates through the configured recipe/training path.\n- **Commit:** evaluate candidates, apply selection policy, and publish accepted artifacts as version history.\n\nUse this model to diagnose where a learning loop is failing instead of changing several layers at once.\n\n## Working Method\n\n1. Read the relevant bundled reference before changing commands or config.\n2. Identify the user's current surface: harness, weights, or test-time training.\n3. Identify the measurable evaluator or feedback signal.\n4. Map the system onto Serve → Observe → Grow → Commit.\n5. Make the smallest change that tests one hypothesis.\n6. Verify with a health check, evaluator result, artifact/version change, or recipe-specific test.\n7. Preserve rollback/version history when proposing changes to a live agent.\n\n## References\n\nStart with the smallest relevant file:\n\n- `references/README.md` — project overview and architecture.\n- `references/intro.rst` — conceptual introduction.\n- `references/installation.rst` — installation requirements.\n- `references/quickstart.rst` — initial end-to-end setup.\n- `references/core-loop.rst` — learning-cycle model.\n- `references/harness-adapters.rst` — adapting agent harnesses.\n- `references/write-a-harness-method.rst` — implementing harness improvement methods.\n- `references/write-a-recipe.rst` — recipe authoring.\n- `references/configuration.rst` — configuration reference.\n- `references/cli.rst` — command-line reference.\n\n## Safety and Accuracy\n\n- Never claim training, evaluation, deployment, or an update succeeded without fresh verification evidence.\n- Do not invent Reef configuration keys or CLI flags.\n- Do not expose credentials, tokens, or private training data.\n- Treat feedback datasets and interaction records as potentially sensitive.\n- Prefer reversible, versioned changes to live-agent artifacts.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}