← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORY

Update to NVIDIA BioNeMo Agent Toolkit

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "drug-discovery-pipeline",
  "description": "Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs: generate drug-like molecules with GenMol, dock them to a protein target with DiffDock, then predict binding affinity with Boltz2. Use this skill whenever the user wants to generate and screen small molecule drug candidates, perform hit discovery, optimize leads against a protein target, or do virtual screening combining molecule generation, docking, and affinity prediction. Triggers on: drug discovery pipeline, hit discovery, lead optimization, virtual screening, molecule generation, molecular docking, binding affinity, GenMol, DiffDock, Boltz2, SMILES, SAFE notation, NIM microservice. This is a multi-step pipeline composing three BioNeMo NIMs.",
  "included_files": [],
  "skill_md_contents": "---\nname: drug-discovery-pipeline\ndescription: >\n  Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs:\n  generate drug-like molecules with GenMol, dock them to a protein target with DiffDock,\n  then predict binding affinity with Boltz2. Use this skill whenever the user wants to\n  generate and screen small molecule drug candidates, perform hit discovery, optimize\n  leads against a protein target, or do virtual screening combining molecule generation,\n  docking, and affinity prediction. Triggers on: drug discovery pipeline, hit discovery,\n  lead optimization, virtual screening, molecule generation, molecular docking, binding\n  affinity, GenMol, DiffDock, Boltz2, SMILES, SAFE notation, NIM microservice. This is\n  a multi-step pipeline composing three BioNeMo NIMs.\nlicense: Apache-2.0 AND CC-BY-4.0\nallowed-tools: Bash, Read, Write, AskUserQuestion\n---\n\n# Drug Discovery Pipeline\n\nScreen drug candidates end-to-end using three BioNeMo NIMs in sequence:\n\n```\nStep 1: GenMol    →  Step 2: DiffDock  →  Step 3: Boltz2\n(Generate mols)      (Dock to target)      (Predict affinity)\n```\n\n---\n\n## Overview\n\nThis pipeline is used for:\n- **De novo hit discovery**: generate drug-like molecules and screen them against a target\n- **Lead optimization**: start from a known scaffold and generate improved analogs, then dock and score\n- **Virtual screening**: dock a library of candidates and filter by docking confidence + affinity\n\n---\n\n## Before you start\n\nConfirm with the user:\n1. **Target protein**: PDB file or sequence of the binding target\n2. **Starting point**: de novo (no scaffold) or scaffold decoration (known core)?\n3. **Scoring**: drug-likeness (QED) or lipophilicity (LogP)?\n4. **API mode**: hosted or local Docker?\n\nFor local Docker, do not assume all NIMs are running on `localhost:8000` at the\nsame time. Either run one container at a time and hand files/results between\nsteps, or start each NIM on a distinct host port and set the per-step URLs.\n\n---\n\n## Step 1: Generate molecules with GenMol\n\nGenMol requires SAFE notation input (not raw SMILES). Use the `safe-mol` package.\n\n```python\nimport requests, json, os\nimport safe as sf                          # pip install safe-mol\nfrom pathlib import Path\n\nNGC_API_KEY = os.environ[\"NGC_API_KEY\"]\nHOSTED = True\n\nif HOSTED:\n    genmol_url = \"https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate\"\n    headers = {\"Content-Type\": \"application/json\",\n               \"Authorization\": f\"Bearer {NGC_API_KEY}\"}\nelse:\n    genmol_url = \"http://localhost:8000/generate\"\n    headers = {\"Content-Type\": \"application/json\"}\n\n# De novo generation (no scaffold):\nsafe_input = \"[*{20-30}]\"\n\n# Scaffold decoration (known core):\n# scaffold_smiles = \"c1ccccc1\"\n# safe_input = sf.encode(scaffold_smiles) + \".[*{5-10}]\"\n\npayload = {\n    \"smiles\": safe_input,              # field is named 'smiles' but takes SAFE notation\n    \"num_molecules\": 30,               # request more to compensate for post-generation filtering\n    \"scoring\": \"QED\",                  # QED or LogP\n    \"unique\": True,\n    \"temperature\": \"1.0\",             # NOTE: must be string, not float\n    \"noise\": \"1.0\",                   # NOTE: must be string, not float\n}\n\nr = requests.post(genmol_url, headers=headers, json=payload)\nr.raise_for_status()\nmolecules = r.json()[\"molecules\"]\nmolecules_sorted = sorted(molecules, key=lambda x: x[\"score\"], reverse=True)\ntop_20 = molecules_sorted[:20]\n\nprint(f\"Generated {len(molecules)} valid molecules (requested 30)\")\nprint(\"Top 5 by QED score:\")\nfor m in top_20[:5]:\n    print(f\"  {m['smiles'][:50]}  score={m['score']:.4f}\")\n```\n\n---\n\n## Step 2: Dock molecules with DiffDock\n\nPrepare the protein and dock each candidate:\n\n```python\n# Load protein (ATOM records only)\nreceptor_pdb_raw = Path(\"target.pdb\").read_text()\nreceptor_pdb = \"\\n\".join(line for line in receptor_pdb_raw.splitlines()\n                          if line.startswith(\"ATOM\"))\n\nif HOSTED:\n    diffdock_url = \"https://health.api.nvidia.com/v1/biology/mit/diffdock\"\nelse:\n    diffdock_url = \"http://localhost:8000/molecular-docking/diffdock/generate\"\n\ndocking_results = []\n\nfor i, mol in enumerate(top_20):\n    payload = {\n        \"protein\": receptor_pdb,\n        \"ligand\": mol[\"smiles\"],\n        \"ligand_file_type\": \"txt\",     # \"txt\" for SMILES input\n        \"num_poses\": 5,\n        \"time_divisions\": 20,\n        \"steps\": 18,\n        \"save_trajectory\": False,\n    }\n\n    r = requests.post(diffdock_url, headers=headers, json=payload)\n    r.raise_for_status()\n    result = r.json()\n\n    best_conf = result[\"position_confidence\"][0]  # rank 1 pose\n    best_pose = result[\"ligand_positions\"][0]\n\n    docking_results.append({\n        \"smiles\": mol[\"smiles\"],\n        \"qed_score\": mol[\"score\"],\n        \"docking_confidence\": best_conf,\n        \"best_pose_sdf\": best_pose,\n    })\n    print(f\"  Mol {i+1:2d}: QED={mol['score']:.3f}  docking_conf={best_conf:.4f}\")\n\n# Rank by docking confidence\ndocking_results.sort(key=lambda x: x[\"docking_confidence\"], reverse=True)\nprint(f\"\\nTop 3 by docking confidence:\")\nfor d in docking_results[:3]:\n    print(f\"  {d['smiles'][:50]}  conf={d['docking_confidence']:.4f}\")\n```\n\n---\n\n## Step 3: Predict binding affinity with Boltz2\n\nFor the top docking candidates, predict structure-based binding affinity:\n\n```python\nif HOSTED:\n    boltz_url = \"https://health.api.nvidia.com/v1/biology/mit/boltz2/predict\"\nelse:\n    boltz_url = \"http://localhost:8000/biology/mit/boltz2/predict\"\n\n# Use the target protein sequence (not PDB)\ntarget_sequence = \"<YOUR_TARGET_PROTEIN_SEQUENCE>\"\n\naffinity_results = []\nfor d in docking_results[:5]:  # score top 5 docking hits\n    payload = {\n        \"polymers\": [\n            {\"id\": \"A\", \"molecule_type\": \"protein\", \"sequence\": target_sequence}\n        ],\n        \"ligands\": [\n            {\"id\": \"L1\", \"smiles\": d[\"smiles\"], \"predict_affinity\": True}\n        ],\n        \"recycling_steps\": 3,\n        \"sampling_steps\": 50,\n        \"diffusion_samples\": 1,\n        \"output_format\": \"mmcif\",\n    }\n\n    r = requests.post(boltz_url, headers=headers, json=payload)\n    r.raise_for_status()\n    result = r.json()\n\n    aff = result[\"affinities\"][\"L1\"]\n    pic50 = aff[\"affinity_pic50\"][0]\n    prob_binding = aff[\"affinity_probability_binary\"][0]\n\n    affinity_results.append({\n        **d,\n        \"pic50\": pic50,\n        \"probability_binding\": prob_binding,\n    })\n    print(f\"  {d['smiles'][:40]}  pIC50={pic50:.2f}  P(bind)={prob_binding:.3f}\")\n\n# Final ranking by pIC50\naffinity_results.sort(key=lambda x: x[\"pic50\"], reverse=True)\n```\n\n---\n\n## Interpreting results\n\n- **GenMol QED score**: 0–1; >0.5 is drug-like\n- **DiffDock confidence**: higher = more reliable binding pose prediction\n- **Boltz2 pIC50**: predicted -log10(IC50); >6 = sub-micromolar, >8 = very potent\n- **P(bind)**: probability of binary binding; >0.7 = likely binder\n\n---\n\n## Quick reference — skill dependencies\n\n| Step | Skill | Key endpoint |\n|---|---|---|\n| Molecule generation | `genmol-nim` | `/biology/nvidia/genmol/generate` |\n| Docking | `diffdock-nim` | `/molecular-docking/diffdock/generate` |\n| Affinity prediction | `boltz2-nim` | `/biology/mit/boltz2/predict` |\n"
}

SHA-256: 095c7dbe4d607b95f5ca7ca28900b3139f5b0ecee26c4b06c274061651e0e9b9