{"id":20842,"plugin_id":"plugins_6aa9e5854d588191a0a6bf357da86329","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:16:35.038Z","digest":"4b05ce0963c9776ceea2f77380040f92ca69001895c73a6ae336f827a91dc14f","against":null,"payload":{"description":"Build and run custom agents with Recurse for tasks that benefit from test-time compute and iterative refinement. Examples include design exploration, performance optimization, and generating artifacts that improve through verification and revision. Use this skill when designing an adaptive component that solves a specific problem. When necessary, clarify uncertain requirements on the goal and constraints.","included_files":[],"name":"recurse","skill_md_contents":"---\r\nname: recurse\r\ndescription: Build and run custom agents with Recurse for tasks that benefit from test-time compute and iterative refinement. Examples include design exploration, performance optimization, and generating artifacts that improve through verification and revision. Use this skill when designing an adaptive component that solves a specific problem. When necessary, clarify uncertain requirements on the goal and constraints.\r\n---\r\n\r\n# Recurse\r\n\r\nRecurse hosts custom agents that work with Python tools. You shape the task, tools, and verifiers to define custom agents, they iterate on candidate results without you supervising their every step. You author the agents locally; `recurse run` executes them on a serverless runtime. Deploying a custom agent as an MCP makes it available for ongoing future use.\r\n\r\n## When to choose Recurse\r\n\r\nRecurse is especially useful when:\r\n- The task is amenable to iterative approaches where feedback can guide successive refinement of a candidate result, or exploration of other candidates.\r\n- There are many such approaches to try.\r\n\r\nThis involves a substantial search over a design space: try variants, tune performance, refine artifact quality. The aim is to discover progressively better approaches (e.g. tools, prompts, verifiers) to improve the objective while meeting constraints. The best achievable result need not be known.\r\n\r\nFor example:\r\n- **Design a lighter mounting bracket.** Given a CAD model, material, and load cases, find a lighter manufacturable design that satisfies strength and stiffness limits. A general-purpose harness can make an edit and run a simulation, but a thorough exploration requires many competing attempts. A Recurse agent varies geometry, uses numerical stress and displacement checks, and preserves the lightest feasible candidate.\r\n- **Generate diverse puzzle levels.** Produce levels matching a concept, mechanic mix, difficulty, and play time while avoiding similarity to a growing corpus. This is a multidimensional search requiring substantial trial and error; general-purpose harnesses tend to follow a generic workflow and drift from the search objectives. A Recurse agent explores variants using solvability verifiers and numerical measures of mechanic usage, solution length, difficulty, and corpus similarity.\r\n- **Generate 3D objects at scale.** A general-purpose harness can create a one-off 3-D model of a bag or a shoe from a picture, but handcrafting the last object-specific details makes scaling impractical. Many objects share common construction and rendering operations, but details like straps, pockets, soles, and laces require special tools. Recurse agents iterate through tool combinations and compare renders. When missing capabilities block progress, you arm the agents with new tools, expanding what future incarnations can produce.\r\n\r\n## The Recurse harness\r\n\r\nA Recurse agent calls an LLM in a special, serverless harness whose FSM structure guides the LLM to iteratively solve problems via Python tools. You focus on defining the prompt, tools, and verifiers for your custom agent, create an `agent.yaml` manifest file to inventory them, and use the Recurse SDK to bundle them for execution in the harness. The harness executes the LLM's tool calls, retains their results, and presents measurements/errors to the LLM for reasoning and decision-making. This lets it explore alternatives and refine results across iterations.\r\n\r\nThe LLM may request several tool calls at the same time, and chain one tool's result into another. The harness automatically resolves such dependencies: calls wait for the values they require, while independent calls can run simultaneously for fast execution. Not all tool composition has to take place over a single iteration: the harness provides a way for the LLM to store tool return values in program memory as Python objects, collect results over multiple rounds, and compose them together when it has all the values it needs. In such cases, the LLM can simply refer to past return values without needing to reproduce them in its responses. To generate artifacts for retrieval after the run, simply provide a tool that writes a file under `recurse.context().workspace`.\r\n\r\nWhen registering tools in the manifest file, you can use the following parameters to configure the harness in order to fine-tune or optimize your tools' interaction with Recurse.\r\n\r\n| Concept | How it works |\r\n| --- | --- |\r\n| `storable` | Controls whether a tool allows the harness to save its return value in program memory. By default, a non-`None` return type enables it. Only disable this for tools that produce nothing but textual results for the LLM context, and/or no Python object that can compose with others. |\r\n| `no_storage` | Parameter names to exclude from Python object passing, forcing the LLM to supply inline values. Use this for identifier or control-flag parameters where Python object storage/passing is unnecessary. Default: empty list. |\r\n| `volatile` | Set to `true` when identical arguments can produce different results, such as sampling or querying a simulator with inherent randomness. The LLM may call such tools with identical arguments, or in recurring patterns, without alerting the harness of a stuck loop. Default: `false`. |\r\n| `pure_args` | Parameter names the tool guarantees not to mutate in place. Recurse uses this to plan/optimize the execution of tool calls. Default (empty list) is conservative (no purity); use ``null`` as a shortcut to indicate \"all parameters are pure\". It is important to declare this correctly, otherwise tools may race and produce incorrect results. This is an optimization flag; when unsure, use the default value. |\r\n\r\n## Start with a proposal\r\n\r\nThe user may not provide full clarity on:\r\n- The input/output contract, artifact formats (if any), or sample inputs/outputs.\r\n- A clear objective/goal to optimize for.\r\n- Verification methods to use\r\n- Any hard constraints\r\n- Time and/or spending allowances\r\n\r\nBefore any cloud work, come up with a concrete proposal that offers clarity on these. In this process, ask about ambiguities that materially affect usefulness; state narrow assumptions for routine details. Translate general intent into constraints, checks and verifiers. Try to use arbitrary weights or acceptance thresholds as sparingly as possible. When unavoidable, present them as proposals with their trade-offs.\r\n\r\nWhen proposing a plan, explain your framing of the problem, how you plan to solve it, and why it fits. If you recommend Recurse, explain what the custom agent will own; describing a generic loop alone leaves that unclear.\r\n\r\n## Agent design and learning from evidence\r\n\r\nKeep tools intentional: actions apply the agent's choices or proposals; independent verifiers return measurements, validation results, and corrective diagnostics; and a termination tool decides whether the task is complete and, if so, saves the final result and an authoritative receipt. Do not leak target answers or weaken acceptance criteria. See [tool guidance](#authoring-tools).\r\n\r\nStart with a baseline approach. Since Recurse allows you to run multiple agents in parallel, always consider which distinct approaches you can explore in parallel. Use concurrent exploration whenever the approaches are independent and your allowance supports it; use sequential attempts when later choices depend on earlier results. Discuss your parallelism strategy in the plan.\r\n\r\nWhen running agents in parallel, make sure to provide the necessary isolation when/if necessary. Evaluate, inspect or score results of candidate agents with the same criteria, make sure not to compare apples with oranges in cases where you co-evolve evaluation criteria with the agents. Choose concurrency to fit the work and resources, not an arbitrary worker count.\r\n\r\nWhen necessary, revisit whether the constraints, the verification tools or your evaluation mechanism still capture the user's intent as you inspect results or encounter new kinds of inputs. For example, when building an object reconstruction agent that creates 3-D models from 2-D images, a check that captures a backpack's shape may overlook a shoe's laces. Use those mismatches to strengthen or generalize the concrete interpretation of the same intent, retaining earlier requirements. Version verifiers, checks and evaluation criteria as you revise them; re-evaluate baseline and contenders consistently; keep old scores as history rather than ranking scores from different evaluators together. Agree materially new requirements with the user before changing what counts as acceptable.\r\n\r\nForm theses (multiple at a time, if you can) on how you improve the agent. Within your parallelism budget, try to validate these theses and see if they result in actual improvement. Prefer small changes to prompt, tools, or verifiers to large re-architectures unless you have data suggesting the latter is necessary. Compare success, quality, trials, time, and cost on the same cases, criteria, and budgets; repeat variable results. Repairing or improving one case is not proof of generalization, you must have a clear reasoning as to how/why each thesis helps. Conversely, do not invalidate a thesis only because of metric/measurable regressions. Invalidating a thesis requires resolving the conflict between your initial reasons for formulating the thesis, and the empirical results. Doing this carefully may help you repair a faulty implementation of an otherwise useful thesis, or else extract the necessary learnings to formulate better theses.\r\n\r\nSeparate feasibility from quality. All checks must pass before a candidate can win on its objective. A high score cannot compensate for an unmet requirement. Preserve the best feasible candidate when later attempts regress. If your agents consistently fail to produce feasible results, scrutinize tools and verifiers.\r\n\r\n### Prompting\r\n\r\nCreate a `prompts.md` file to house your custom agent's instructions for the specific problem it will solve, and set `agent.prompt: prompts.md` in `agent.yaml`. Use the final proposal, the actual tools, and verifiers as a starting point and/or a reference. Explain the following clearly:\r\n- What outcome to improve, how to measure it, and what makes one feasible result better than another.\r\n- Which requirements are hard constraints, what degrees of freedom the agent can explore, and which domain facts constitute implicit constraints.\r\n- What the verifiers measure, how those measurements relate to the overall goal and constraints, and what they do not establish. Include domain knowledge necessary to interpret failures and trade-offs.\r\n- What defines completion or convergence for the agent; i.e. when does the agent stop iterative refinement and report results.\r\n- What evidence and artifacts the agent must return; and any conditions, allowances or restrictions that the agent must comply with.\r\n\r\nFor the mounting-bracket example, you would explain that the agent is searching for lower mass while meeting strength, stiffness, and manufacturing requirements under the given material and load constraints. You would describe how the available checks and verifiers establish those requirements. You would instruct the agent to explore geometric changes and competing design patterns (maybe along with certain tactics for doing so); but not prescribe a sequence of CAD edits.\r\n\r\nEncourage the agent to develop hypotheses, explore alternatives, accept/reject hypotheses with the methodology in your instructions above, and use results to decide which avenues merit further work. Give it enough domain context to make those decisions independently. Do not turn the prompt into an ordinary workflow without iterative refinement unless the problem is very simple. Do not dictate one tool call per turn, or supply a sequence of candidate answers.\r\n\r\nTool signatures and parameter instructions belong in tool docstrings (see [tool guidance](#authoring-tools)); the prompt explains how the tools support this problem's objective.\r\n\r\nIn certain classes of problems, you may want to use \"control\" inputs and vary them during agent design. Use the `{{ input.NAME }}` substitution mechanism to interpolate control inputs into `inputs.task`. Use this facility to experiment with control knobs, vary the instructions in `prompts.md` to make \"architectural\" changes.\r\n\r\nImprove the prompt from the runs you observe. If agents consistently misinterpret a measurement, ignore a constraint, or explore an unproductive part of the search space, identify the missing or misleading instruction. Revise the prompt and compare with the baseline on the same cases, verifiers, and allowances. A longer prompt or a successful single run does not establish improvement on its own.\r\n\r\n### Authoring tools\r\n\r\nDefine tools as top-level Python functions in your application's tool file, usually `tools.py`. Each function's name becomes its tool name. Prefix helpers with `_`.\r\n\r\nThe harness uses each function's Google-style docstring to describe the tool to the LLM. The summary, body, and `Returns:` section become the tool description; `Args:` entries become parameter\r\ndescriptions, and type annotations supply the schema. Write these as instructions the LLM will use to choose and call the tool: explain its purpose, constraints, effects, inputs and its return value.\r\n\r\nYou must fully annotate parameters and returns. When using generic types (e.g. `list`), supply the type parameter(s) whenever possible; the harness reflects this information in the JSON schema of the tool to better guide the LLM. Do not use bare generics unless there really are no constraints on the type parameter(s), or they are really unknown.\r\n\r\nInclude an `Args:` entry for every parameter and a `Returns:` description for non-`None` results, including units or valid ranges. Omit `Returns:` for `-> None`. Use safe defaults and raise errors that report not only what was wrong, but also what the agent can correct. Keep tools to the point and safe to retry where possible. Never use mutable global variables to share \"hidden state\" across tools. The harness cannot see such dependencies, and this may result in incorrect execution schedules.\r\n\r\nFor example, the following two tools demonstrate construction, an independent measurement, and a state object. The measurement is illustrative; but the style is suggestive. Use the actual domain validators for your application.\r\n```python\r\nfrom dataclasses import dataclass\r\n\r\ntype Point = tuple[float, float]\r\n\"\"\"A 2-D point.\"\"\"\r\n\r\n\r\ndef _cross(o: Point, a: Point, b: Point) -> float:\r\n    \"\"\"Compute the z-component of the cross product of ``(a - o)`` and ``(b - o)``.\r\n\r\n    A positive result indicates a left turn, a negative result a right turn, and zero indicates collinearity.\r\n\r\n    Args:\r\n        o: Reference point to calculate the cross product against.\r\n        a: The first term of the cross product.\r\n        b: The second term of the cross product.\r\n\r\n    Returns:\r\n        The z-component of the cross product.\r\n    \"\"\"\r\n    return (a[0] - o[0]) * (b[1] - o[1]) - (a[1] - o[1]) * (b[0] - o[0])\r\n\r\n\r\ndef _on_segment(p: Point, a: Point, b: Point) -> bool:\r\n    \"\"\"Whether ``p`` lies on the closed segment between ``a`` and ``b``.\r\n\r\n    Args:\r\n        p: The query point.\r\n        a: The first endpoint of the segment.\r\n        b: The last endpoint of the segment.\r\n    \"\"\"\r\n    return (\r\n        _cross(a, b, p) == 0\r\n        and min(a[0], b[0]) <= p[0] <= max(a[0], b[0])\r\n        and min(a[1], b[1]) <= p[1] <= max(a[1], b[1])\r\n    )\r\n\r\n\r\n@dataclass(frozen=True)\r\nclass Polygon:\r\n    \"\"\"A simple, non-self-intersecting polygon with at least three vertices.\"\"\"\r\n\r\n    points: tuple[Point, ...]\r\n    \"\"\"Points on the periphery of the polygon.\"\"\"\r\n\r\n    def __post_init__(self) -> None:\r\n        pts = self.points\r\n        n = len(pts)\r\n        if n < 3:\r\n            raise ValueError(\"Cannot define a polygon with fewer than three points\")\r\n\r\n        edges = [(pts[i], pts[(i + 1) % n]) for i in range(n)]\r\n\r\n        # 1. No vertex may lie on an edge it isn't an endpoint of.\r\n        for k, p in enumerate(pts):\r\n            for i, (a, b) in enumerate(edges):\r\n                if i not in (k, (k - 1) % n) and _on_segment(p, a, b):\r\n                    raise ValueError(f\"Vertex {k} lies on edge {i}\")\r\n\r\n        # 2. No two edges may properly cross.\r\n        for i in range(n):\r\n            for j in range(i + 1, n):\r\n                (p1, p2), (q1, q2) = edges[i], edges[j]\r\n                if (\r\n                    _cross(q1, q2, p1) * _cross(q1, q2, p2) < 0\r\n                    and _cross(p1, p2, q1) * _cross(p1, p2, q2) < 0\r\n                ):\r\n                    raise ValueError(f\"Edges {i} and {j} cross\")\r\n\r\n\r\ndef build_polygon(points: list[Point]) -> Polygon:\r\n    \"\"\"Construct a non-self-intersecting polygon from the points on its periphery.\r\n\r\n    Args:\r\n        points: Points on the periphery of the polygon. Each point connects with the next via an edge. The last point in the list connects with the first point. The points must define a non-self-intersecting polygon.\r\n\r\n    Returns:\r\n        The ``Polygon`` the given list of points represents.\r\n\r\n    Raises:\r\n        ValueError: If fewer than three points are given, or if the points define a self-intersecting polygon.\r\n    \"\"\"\r\n    return Polygon(tuple(points))\r\n\r\n\r\ndef measure_area(polygon: Polygon) -> float:\r\n    \"\"\"Measure the given polygon's area.\r\n\r\n    Args:\r\n        polygon: The polygon whose area to measure.\r\n\r\n    Returns:\r\n        Area of the given polygon.\r\n    \"\"\"\r\n    vertices = polygon.points\r\n    n = len(vertices)\r\n    area = 0.0\r\n\r\n    # Shoelace formula: sum the cross products of consecutive vertices.\r\n    for i in range(n):\r\n        x_i, y_i = vertices[i]\r\n        x_j, y_j = vertices[(i + 1) % n]\r\n        area += x_i * y_j - x_j * y_i\r\n\r\n    return abs(area) / 2.0\r\n```\r\n\r\nIn `agent.yaml`, set `tools.source` to the application-relative path of the tool file. Register every public function implementing a tool in that file by its exact name in `tools.register`:\r\n```yaml\r\ntools:\r\n  source: tools.py\r\n  register:\r\n    build_polygon: ~\r\n    measure_area: ~\r\n```\r\n\r\n`~` uses the defaults. Set per-tool `volatile` or `storable` individually when necessary; use `tools.defaults` for settings common to all tools. The `tools.built_in` flag (default: `true`) enables/disables built-in tools like note-taking, TODO management and planning, and others.\r\n\r\n#### Designing tools that compose\r\n\r\nReturn values that other tools can accept directly. In the example, `build_polygon` returns a `Polygon` and `measure_area` accepts a `Polygon`, enabling the agent to chain them without reconstructing the candidate. The agent chooses whether and when to use that chain.\r\n\r\nCarry results and evolving state through parameters with intentional types and return values. Returning only a prose report or transmitting results through mutable globals hides the relationship between tools. Properly describe the return value (and its uses, if not obvious) in the docstring, so the agent can recognize useful combinations.\r\n\r\n#### Actions, validators, and results\r\n\r\nAction tools apply the agent's choices: construct candidates, edit them, or take a certain step to interact with some external system. They may enforce syntax and domain invariants to facilitate failing fast. Validator tools use independent verifiers (e.g., tests, scorers, simulators, solvers, queries) to gauge feasibility or outcome quality. They return observable measurements and concise diagnostics that can influence the next action. Independence of validators from action tools is very important; the former encodes *what* a good outcome is, the latter encodes *how* to make progress towards it. Validators must not tautologically rubber-stamp results.\r\n\r\nGive the agent one explicit completion tool that decides whether the task is complete, and if it is, returns the best candidate along with an authoritative receipt. Register it alongside the action and validator tools. Use artifacts for durable evidence and larger files; do not make an artifact the only explanation of the result. The agent's final JSON must match `outputs` in `agent.yaml`, and agree with the receipt.\r\n\r\nInside tools, `recurse.context().inputs` is a read-only mapping of caller inputs (excluding the `inputs.task`, which becomes part of the prompt). Write artifact files under `recurse.context().workspace`.\r\n\r\n### Never do these\r\n\r\n- Narrow the agent into an ordinary workflow that decides every action in advance unless the problem is very simple and solvable through such an approach.\r\n- Limit the agent to executing one tool at a time regardless of actual dependencies. The harness handles tool execution schedules automatically.\r\n- Use mutable global variables to transmit tool results or evolving run state. Doing this may result in tool races.\r\n- Accept a candidate's quality with no independent validation.\r\n- Weaken constraints or acceptance criteria to manufacture success.\r\n- Compare candidates with incompatible evaluators or allowances.\r\n- Overwrite/lose the best feasible candidate while exploring worse attempts.\r\n- Claim an artifact, successful result, or global optimum without supporting evidence.\r\n\r\n## When to finalize an agent design\r\n\r\nUse a real acceptance target if the user provides one. Otherwise, judge diminishing returns by tracking gains, remaining approaches, and the cost of another attempt. Explain/justify that judgment using the history of runs; an arbitrary number of low-gain attempts alone does not establish diminishing returns. Do not invent arbitrary score thresholds or claim global optimality unless you have strong evidence for it. Track aggregate usage across runs and retain hard limits. Do not start work that cannot fit the remaining allowance.\r\n\r\nBefore normal completion, save the best agent design and return its measurable facts. State whether you are stopping due to acceptance, diminishing returns, a resource limit, or a blocker. A runtime timeout is not evidence of diminishing returns. After an abrupt failure, report only the evidence you have; do not assume that an artifact exists if you cannot locate it. Diagnostic artifacts are not final results.\r\n\r\n## Package the agent\r\n\r\n```text\r\nmy-agent\r\n├── agent.yaml      # Identity, runtime, prompt path, input/output schemas, tools\r\n├── prompts.md      # Instructions for the custom agent\r\n├── tools.py        # Python tools with type annotations and their supporting classes\r\n├── pyproject.toml  # Dependencies and explicit package build backend\r\n└── uv.lock         # Exact dependency lockfile generated by uv\r\n```\r\n\r\nUse Python 3.14 or newer. Install the CLI with `pip install recurse-sdk`. The application needs `agent.yaml`, `pyproject.toml` with an explicit build backend, `uv.lock`, and its prompt and tools. Configure the backend to include the exact lockfile and all runtime files in the source distribution; exclude `.env`, caches, and virtual environments. Add `recurse-sdk` to the application dependencies if tools import `recurse`.\r\n\r\nSelect a model in `agent.model`, such as `gpt-6-astra`, or omit it for the service default, currently `gpt-5.6-luna`. Preparation records that selection with the version; unsupported models fail. There is no model CLI flag; deploy a new version to change the model. Compare quality and cost before choosing a more expensive model.\r\n\r\n## Run or reuse\r\n\r\nUse `run` for a one-off execution or while shaping/designing new agents. Deploy as MCP after evidence shows the agent is reusable.\r\n\r\n```sh\r\nrecurse login\r\nrecurse run ./my-agent --inputs inputs.json\r\nrecurse deploy ./my-agent --as mcp\r\n```\r\n\r\nRuns have a 15-minute execution limit. Use `recurse status <run-id>` to inspect the run or `recurse cancel <run-id>` to request cancellation. Download available artifacts within 24 hours using `recurse artifacts <run-id> --output results`.\r\n\r\nConnect an MCP host via `recurse mcp serve <deployment-id>`. For Codex, set `startup_timeout_sec = 180` and `tool_timeout_sec = 1140` in the server configuration so a long run is not cut short by the host. See the [deployment guide](https://recurse.run/docs/deploy).\r\n\r\n## Troubleshooting\r\n\r\nIf you're experiencing a persistent issue with Recurse, try these steps:\r\n- Upgrade `recurse-sdk`.\r\n- Refresh the skill.\r\n- Check the relevant section in the [documentation](https://recurse.run/docs/).\r\n\r\n## Using the Recurse CLI\r\n\r\n### Install and authenticate\r\n\r\n```sh\r\npip install recurse-sdk\r\nrecurse login\r\n```\r\n\r\nLogin opens the browser and stores the device credential in the operating system keychain. `recurse logout` revokes this device's login. Keep credentials out of the application and MCP configuration. Use `recurse --help` or `recurse <command> --help` to inspect the installed version.\r\n\r\n### Run an application\r\n\r\n```sh\r\nrecurse run ./my-agent --inputs inputs.json\r\n```\r\n\r\n`inputs.json` must contain one JSON object matching the application's input schema. Omitting `--inputs` supplies `{}`. The CLI applies schema defaults and validates values. To read from stdin and set resource ceilings, use:\r\n\r\n```sh\r\nrecurse run ./my-agent --inputs - --cpu 2 --memory-mib 2048 < inputs.json\r\n```\r\n\r\nBoth `run` and `deploy` default to 1 CPU and 1024 MiB. `--cpu` accepts 0.125–16 in 0.125 increments; `--memory-mib` accepts 512–16384 in 128 MiB increments. These are billable ceilings; select them to fit requirements of the problem and your allowance.\r\n\r\nThe CLI packages, uploads, and prepares the application, prints `run: <run-id>`, and waits for its terminal state. It then prints `status`, the result or error when present, and the artifact count. Runs have a 15-minute execution limit. A run ID identifies one execution; running the command again starts another execution.\r\n\r\n| Direct-run exit code | Meaning |\r\n| --- | --- |\r\n| `0` | Run is successful; inspect the receipt to evaluate the results. |\r\n| `1` | Confirmed agent failure. |\r\n| `2` | CLI-layer error. |\r\n| `3` | Confirmed remote cancellation. |\r\n| `4` | Confirmed infrastructure failure. |\r\n| `5` | Confirmed agent timeout. |\r\n| `130` | CLI interrupted by Ctrl-C; inspect the reported remote state. |\r\n\r\nCodes `1`, `3`, `4`, and `5` are confirmed terminal run states. Code `2` covers every failure produced by the CLI rather than the run: invalid command syntax or option values, local input and packaging errors, authentication failures, request and transport failures, malformed service responses, and `observation_timeout` where remote state is unconfirmed. On `2`, inspect stderr and any printed run state before deciding whether to retry.\r\n\r\n| Public error | Next step |\r\n| --- | --- |\r\n| `insufficient_balance` | Check `recurse billing balance`; obtain user approval before adding funds. |\r\n| `secret_unavailable` | Check `recurse secret list` and the required binding. |\r\n| `invalid_inputs` | Correct `--inputs` against the input schema in `agent.yaml`. |\r\n| `invalid_agent` | Check the declaration and packaged tools. |\r\n| `invalid_output` | Check the returned result against the declared output schema. |\r\n| `artifact_failed` | Check artifact paths and retain the run ID. |\r\n| `execution_failed`, `infrastructure_failed`, `unknown_error` | Retain the run ID and available evidence; do not invent a cause. |\r\n\r\n`authentication_failed` directs you to `recurse login`. `request_failed` and `observation_timeout` do not establish that remote execution stopped. If state is unknown, execution and charges may continue: inspect the existing run before rerunning. Keep the admission reference if no run ID was obtained.\r\n\r\n### Inspect, cancel, and retrieve\r\n\r\nCtrl-C during an admitted `recurse run` requests cancellation and exits `130`, even if the run finishes first. Pending or unconfirmed cancellation means execution and charges may continue; inspect the reported state. A second Ctrl-C stops waiting. Use the printed run ID:\r\n\r\n```sh\r\nrecurse status <run-id>\r\nrecurse cancel <run-id>\r\nrecurse artifacts <run-id> --output results\r\n```\r\n\r\nOn POSIX terminals, Ctrl-Z suspends the local CLI; `fg` or `bg` resumes it without cancelling remote work.\r\n\r\n`status` reports the current state and available result/error and artifact count. Inspect again\r\nafter requesting cancellation to see the terminal state. Download available artifacts within\r\n24 hours after completion. Use a new destination if files already exist: artifact retrieval refuses\r\nto overwrite them. Missing or expired artifacts cannot be reconstructed from a storage key.\r\nAfter failure, inspect preserved evidence before deciding whether a new run is useful.\r\n\r\n### Runtime secrets\r\n\r\nFor a tool that needs a simulator credential:\r\n\r\n```sh\r\nrecurse secret set simulator\r\nrecurse secret list\r\nrecurse run ./my-agent --inputs inputs.json --secret SIM_TOKEN=simulator\r\n```\r\n\r\n`secret set` uses hidden prompts and can create or rotate a secret. `--from-stdin` reads its exact\r\nUTF-8 value from stdin when supplied by a trusted secret source. `--secret ENV=NAME` binds an account\r\nsecret to an environment variable inside tools; repeat it for additional bindings on `run` or\r\n`deploy`. The value itself must not appear in the manifest, command arguments, or application archive.\r\nRotation affects future runs; active runs retain their selected secret version.\r\n\r\n`recurse secret delete simulator` asks for confirmation, destroys the secret, and disables affected\r\ndeployments. Use it only when removal is intended.\r\n\r\n### Deploy and connect through MCP\r\n\r\nUse `run` while shaping an agent. Deploy after evidence shows it is reusable:\r\n\r\n```sh\r\nrecurse deploy ./my-agent --as mcp --cpu 2 --memory-mib 2048\r\nrecurse mcp serve <deployment-id>\r\n```\r\n\r\nAdd `--secret SIM_TOKEN=simulator` to `deploy` if the application needs that binding. Deployment\r\nprepares an immutable application version and prints `deployment`, `endpoint`, and resource defaults.\r\nUse the deployment ID—not a run ID—with `mcp serve`. Configure the MCP host to launch command\r\n`recurse` with arguments `[\"mcp\", \"serve\", \"<deployment-id>\"]`; the process serves the connection\r\nthrough local standard input/output using the existing login.\r\n\r\nFor Codex, set `startup_timeout_sec = 180` and `tool_timeout_sec = 1140` in that server's\r\nconfiguration. Give other hosts comparable headroom around the execution limit. See the\r\n[deployment guide](https://recurse.run/docs/deploy) for host setup.\r\n\r\nEditing local files does not update an existing deployment. Run `deploy` again after prompt, tool,\r\ndependency, or model changes and connect the host to the new deployment ID.\r\n\r\n### Balance and credits\r\n\r\n`recurse billing balance` shows total and available balance. `recurse billing redeem CODE` redeems\r\na credit code. `recurse billing top-up 5` opens checkout for $5; supported amounts are $5–$500.\r\nAdd `--no-open` to print the checkout URL instead. Adding funds is a separate purchase decision,\r\nnot an automatic response to an insufficient-balance error.\r\n\r\nSee [documentation](https://recurse.run/docs) and [tool guidance](#authoring-tools).\r\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}