{"id":15480,"plugin_id":"plugin_asdk_app_6aa6b55752908191a29dc03e7b91c419","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:11:28.779Z","digest":"fd56524c0992431f8cacc770507312ea6e7c205c4e6ed9a8e38ccd4f7a95a129","against":null,"payload":{"description":"Compare named assets side by side on a common set of dimensions and return a structured benchmark with explicit evidence strength. Use for direct comparisons, for example 'how does X compare to the other IL-23s', 'benchmark these four programs'. Use identify-analogs when precedents to reason from are wanted rather than a head-to-head table.","included_files":[{"relative_path":"LICENSE","size_in_bytes":802},{"relative_path":"agents/openai.yaml","size_in_bytes":331},{"relative_path":"references/evidence-research.md","size_in_bytes":4449}],"name":"benchmark-assets","skill_md_contents":"---\nname: benchmark-assets\ndescription: \"Compare named assets side by side on a common set of dimensions and return a structured benchmark with explicit evidence strength. Use for direct comparisons, for example 'how does X compare to the other IL-23s', 'benchmark these four programs'. Use identify-analogs when precedents to reason from are wanted rather than a head-to-head table.\"\n---\n\n# Benchmark Assets\n\n## Using this skill\n\nUse the connected Maven Bio MCP server at `https://mcp.mavenbio.com/`. Follow the user's explicit scope, depth, and output preferences; the workflow and output structure below are defaults. Report coverage limits instead of silently narrowing an explicitly requested set.\n\nHyphenated primitive names refer to other skills in this Maven Bio bundle. Consult the relevant skill when composing its workflow. Use the available MCP tool schemas for arguments; pass document identifiers to `read_document` through `ids`, and include a claim-specific `query` when using `format=\"citations\"`.\n\nThis primitive standardizes comparisons across assets without imposing a final table or slide format.\n\n## Use When\n\n- the user wants side-by-side comparison\n- a workflow needs comparable dimensions across multiple assets\n- you need a reusable comparison object before synthesis\n\n## Core Tools\n\n- `match_entity`\n- `search_entities`\n- `research_entity`\n- `search_documents`\n- `read_document` to read a located document (when a benchmark dimension is approval status, designations, or label content, find the label first with `search_documents(source_types=[\"fda_filings\"])` -- `source_types` is a `search_documents` filter, not a `read_document` one)\n\nUse `match_entity` to canonicalize each named asset before benchmarking. Keep the raw asset name in `name` and put sponsor/company/disambiguating text in `context`. Use `search_entities` when you are discovering the benchmark set by criteria rather than starting from a fixed named list.\n\n## Output Contract\n\nReturn benchmark items where each asset has dimensions such as:\n\n- mechanism\n- lead indication\n- highest stage\n- differentiating evidence\n- key risks\n\nEach dimension should include:\n\n- `value`\n- `evidence_rating`\n- `citations`\n- optional `notes`\n\n## Scaling the Benchmark\n\nThe `research_entity`-per-asset path is the right one for small-N comparisons, where\nthe user needs deep narrative context per asset rather than just structured dimensions.\n\nFor larger sets, do not silently degrade the depth of each row to fit the budget.\nNarrow the benchmark set first (tighten the phase, indication, or modality scope via\n`search_entities`), or reduce the number of comparison dimensions, and say which of\nthe two you did. A benchmark of 30 assets at one line each is less useful than a\nbenchmark of 8 assets that actually supports its claims.\n- For independent evidence workstreams, follow the [evidence research procedure](references/evidence-research.md) for each scoped pass, then reconcile the claims before synthesis. Run passes sequentially, or in parallel when the host supports it and the task authorizes it.\n\n## Optional Primitives\n\n- `validate-target` (when a benchmark dimension is target genetic validation or tractability)\n\n## Quality Bar\n\n- keep dimensions comparable across assets\n- separate directly supported facts from analytic judgment\n- when one asset has missing evidence, preserve the asymmetry rather than smoothing it over\n- when benchmarking dimensions touch label content, locate each asset's label with `search_documents(source_types=[\"fda_filings\"])` and read the specific document with `read_document`; prefer `format=\"sections\"` to find the relevant section before pulling full text, and cite the document plus its date on each cell\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}