← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to NVIDIA BioNeMo Agent Toolkit
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "nvmolkit-usage",
"description": "Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF optimization, TFD, conformer RMSD, Butina clustering, and substructure search. Use when the user is importing `nvmolkit.*`, debugging an `nvmolkit` call, choosing between nvMolKit and RDKit for a batched cheminformatics workflow, or wiring nvMolKit results into a torch/numpy pipeline. Out of scope: building nvMolKit from source.",
"included_files": [],
"skill_md_contents": "---\nname: nvmolkit-usage\ndescription: >-\n Write code that calls the installed nvMolKit Python API for GPU-accelerated,\n batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine\n similarity, ETKDG conformer embedding, MMFF/UFF optimization, TFD, conformer\n RMSD, Butina clustering, and substructure search. Use when the user is\n importing `nvmolkit.*`, debugging an `nvmolkit` call, choosing between\n nvMolKit and RDKit for a batched cheminformatics workflow, or wiring nvMolKit\n results into a torch/numpy pipeline. Out of scope: building nvMolKit from\n source.\nlicense: Apache-2.0\nmetadata:\n owner: Kevin Boyd (@scal444)\n risk_tier: skill\n---\n\n# nvMolKit usage\n\n## What nvMolKit is\n\nGPU-accelerated, batched implementations of common RDKit operations. APIs mirror RDKit where possible but are batch-oriented: they take lists of `rdkit.Chem.Mol` (or lists of fingerprints) and process them in parallel on one or more GPUs. nvMolKit links against RDKit at build time; inputs and outputs are real RDKit `Mol` objects.\n\n## Where nvMolKit does well\n\nReach for nvMolKit when:\n\n- The workload is **a large batch of molecules** processed together (typically thousands or more).\n- The metric is **throughput / total wall time across the batch**, not per-molecule latency.\n- The same operation is **repeated identically** across the batch (fingerprinting a library, embedding/minimizing many conformers, bulk pairwise similarity), so the GPU stays saturated.\n\nPlain RDKit is usually the better choice for single-molecule one-offs or workflows that can't be expressed as a batch. nvMolKit is not meant to replace RDKit for those cases.\n\n## Runtime requirements\n\n- An NVIDIA GPU with compute capability 7.0 (V100) or higher\n- A CUDA driver compatible with CUDA 12.6+.\n- A working `torch` install with CUDA support (nvMolKit returns GPU tensors via `torch`'s CUDA array interface).\n\nIf CUDA is unavailable, nvMolKit calls raise. There is no CPU fallback - if the user needs one, use RDKit directly for that path.\n\nWhen helping with installation, make the user choose a PyTorch CUDA backend that the host driver supports before installing nvMolKit. nvMolKit's PyPI wheels are built with CUDA Toolkit 12.9 and depend on CUDA 12 runtime packages, but pip/uv can still select a CUDA 13 PyTorch wheel unless the install command says otherwise.\n\n- Conda: prefer conda-forge `pytorch-gpu`; pin `cuda-version=12.6` or another CUDA version supported by the driver.\n- pip: send the user to the PyTorch install selector (`https://pytorch.org/get-started/locally/`) or previous-versions page (`https://pytorch.org/get-started/previous-versions/`) to install `torch` for a CUDA 12.x backend before installing nvMolKit.\n- uv: install nvMolKit with an explicit backend, e.g. `uv pip install --torch-backend=cu128 nvmolkit`.\n\n## Verify the install before writing real code\n\nRun this once to confirm nvMolKit is importable and a GPU op works end to end:\n\n```python\nimport nvmolkit\nimport torch\nfrom rdkit import Chem\nfrom nvmolkit.fingerprints import MorganFingerprintGenerator\n\nprint(\"nvmolkit:\", nvmolkit.__version__)\nprint(\"cuda available:\", torch.cuda.is_available())\nprint(\"device count:\", torch.cuda.device_count())\n\nmols = [Chem.MolFromSmiles(smi) for smi in [\"CCO\", \"c1ccccc1\", \"CC(=O)O\"]]\nfpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)\nresult = fpgen.GetFingerprints(mols)\ntorch.cuda.synchronize()\nfps = result.torch()\nprint(\"fps shape:\", tuple(fps.shape), \"dtype:\", fps.dtype)\n# Expected: shape (3, 32), dtype torch.int32 (1024 bits packed into 32 int32s per row)\n```\n\nIf this fails, point the user at the install guide on the docs site rather than guessing - see \"Going deeper\" below.\n\n## Entry points\n\n| Task | Module | Primary entry point |\n|---|---|---|\n| Morgan fingerprints | `nvmolkit.fingerprints` | `MorganFingerprintGenerator(radius, fpSize).GetFingerprints(mols)` |\n| Bulk Tanimoto / cosine similarity | `nvmolkit.similarity` | `crossTanimotoSimilarity(...)`, `crossCosineSimilarity(...)`, plus `*MemoryConstrained` variants for results too large to fit in GPU memory |\n| ETKDG conformer embedding | `nvmolkit.embedMolecules` | `EmbedMolecules(molecules, params, confsPerMolecule, ...)` |\n| MMFF94 optimization (one-shot) | `nvmolkit.mmffOptimization` | `MMFFOptimizeMoleculesConfs(molecules, ...)` |\n| UFF optimization (one-shot) | `nvmolkit.uffOptimization` | `UFFOptimizeMoleculesConfs(molecules, ...)` |\n| Forcefield with custom options + constraints | `nvmolkit.batchedForcefield` | `MMFFBatchedForcefield(mols, properties=..., nonBondedThreshold=..., ignoreInterfragInteractions=..., hardwareOptions=...)`, `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`. Per-molecule view `ff[i]` exposes `add_distance_constraint`, `add_position_constraint`, `add_angle_constraint`, `add_torsion_constraint`. Methods: `.compute_energy()`, `.compute_gradients()`, `.minimize(maxIters, forceTol)` |\n| Pairwise conformer RMSD | `nvmolkit.conformerRmsd` | `GetConformerRMSMatrix(mol)`, `GetConformerRMSMatrixBatch(mols)` |\n| Torsion Fingerprint Deviation (TFD) | `nvmolkit.tfd` | `GetTFDMatrix(mol)`, `GetTFDMatrices(mols)` |\n| Butina clustering | `nvmolkit.clustering` | `butina(distance_matrix, cutoff)` (precomputed matrix), `fused_butina(fingerprints, cutoff)` (memory-efficient, on-the-fly) |\n| Substructure search | `nvmolkit.substructure` | `hasSubstructMatch`, `countSubstructMatches`, `getSubstructMatches` |\n| Hardware tuning (batch size, GPU IDs) | `nvmolkit.types` | `HardwareOptions(...)` passed to ETKDG / MMFF / UFF |\n| Optional autotuning of `HardwareOptions` | `nvmolkit.autotune` | `tune_embed_molecules`, `tune_mmff_optimize`, `tune_uff_optimize`, `tune_batched_forcefield`. Requires the `optuna` package |\n\n## Result types and execution model\n\nTwo return shapes carry GPU-resident output, depending on what the operation produces.\n\n### `AsyncGpuResult`\n\nUsed by operations that return a single flat tensor (fingerprints, similarity matrices, RMSD/TFD vectors, Butina inputs). Key behaviors:\n\n- Asynchronous. The kernel may not have completed when the call returns.\n- `result.torch()` returns a zero-copy `torch.Tensor` on the GPU. Caller is responsible for synchronizing before reading values on the host.\n- `result.numpy()` synchronizes and returns a CPU numpy array.\n- Exposes `__cuda_array_interface__`, so it can be passed directly into other nvMolKit functions (e.g. fingerprints → similarity) with no host round-trip.\n\n#### CUDA stream control\n\nA subset of the `AsyncGpuResult`-returning APIs accept an optional `stream: torch.cuda.Stream | None = None` argument so callers can submit nvMolKit work to a non-default stream and overlap it with their own kernels. When omitted, the call uses the current torch stream.\n\nAPIs that take a `stream` argument:\n\n- `MorganFingerprintGenerator.GetFingerprints`\n- `crossTanimotoSimilarity`, `crossCosineSimilarity`, and their `*MemoryConstrained` variants\n- `butina`, `fused_butina`\n- `GetConformerRMSMatrix`, `GetConformerRMSMatrixBatch`\n\nOther APIs (ETKDG, MMFF/UFF optimization, TFD, substructure search) are synchronous to the caller — no stream plumbing needed.\n\nTypical pattern:\n\n```python\nimport torch\nfrom rdkit import Chem\nfrom nvmolkit.fingerprints import MorganFingerprintGenerator\nfrom nvmolkit.similarity import crossTanimotoSimilarity\n\nstream = torch.cuda.Stream()\nfpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)\nmols = [Chem.MolFromSmiles(smi) for smi in [\"CCO\", \"c1ccccc1\", \"CC(=O)O\"]]\n\nwith torch.cuda.stream(stream):\n fps = fpgen.GetFingerprints(mols, stream=stream)\n sim = crossTanimotoSimilarity(fps, stream=stream)\nstream.synchronize()\nprint(sim.torch())\n```\n\n### `Device3DResult`\n\nUsed by ETKDG embedding and MMFF/UFF optimization (one-shot and `BatchedForcefield`) when called with `output=CoordinateOutput.DEVICE`. The GPU-resident equivalent of writing conformers back to `Mol` objects. Fields:\n\n- `values`: `AsyncGpuResult` of shape `(total_atoms, 3)` float64. Concatenated conformer coordinates in CSR-style layout.\n- `atom_starts`, `mol_indices`, `conf_indices`: `AsyncGpuResult` int32 buffers describing the layout (`values[atom_starts[i]:atom_starts[i+1]]` is conformer `i`'s atoms).\n- `energies`, `converged`: `AsyncGpuResult` buffers populated only for MMFF/UFF minimization (not for plain ETKDG).\n- `gpu_id`: device the buffers live on. The `targetGpu` argument on each API picks this; `targetGpu=-1` uses the default consolidation device.\n- `.per_molecule()` returns nested `list[list[torch.Tensor]]` of per-conformer views; `.dense(pad_value=nan)` materializes a padded `(n_mols, max_confs, max_atoms, 3)` tensor.\n\nThe default mode (`CoordinateOutput.RDKIT_CONFORMERS`) still writes optimized coordinates back into each `Mol` and returns Python lists of energies/convergence flags. Reach for `CoordinateOutput.DEVICE` when chaining downstream GPU work (e.g. ETKDG → MMFF → similarity scoring) without host round-trips.\n\n## Configuration\n\nTwo configuration objects expose the GPU/CPU knobs.\n\n### `HardwareOptions` (ETKDG, MMFF, UFF)\n\n`from nvmolkit.types import HardwareOptions`. Passed via `hardwareOptions=` to `EmbedMolecules`, `MMFFOptimizeMoleculesConfs`, `UFFOptimizeMoleculesConfs`, and the `BatchedForcefield` constructors. Every field has an \"auto\" sentinel; the defaults are usually fine.\n\n| Field | Type | Default | Meaning |\n|---|---|---|---|\n| `preprocessingThreads` | int | `-1` (all visible CPUs) | CPU threads for preprocessing |\n| `batchSize` | int | `-1` (auto-tuned) | Number of conformers per GPU batch |\n| `batchesPerGpu` | int | `-1` (auto) | Concurrent batches per GPU; must be `>0` or `-1` |\n| `gpuIds` | `list[int]` | `[]` (all visible GPUs) | Specific device ordinals to target |\n\nPassing a `gpuIds` entry for a device that isn't visible raises `RuntimeError: invalid device ordinal`. For finding good values automatically across a representative sample, see `nvmolkit.autotune` (requires the `optuna` extra); each `tune_*` function returns a `TuneResult` whose `best_config` is a fully-populated `HardwareOptions` ready to pass back into the real call.\n\n`HardwareOptions` round-trips through `to_dict()` / `from_dict()` for persisting tuned configs to disk.\n\n### `SubstructSearchConfig` (substructure search)\n\n`from nvmolkit.substructure import SubstructSearchConfig`. Passed via `config=` to `hasSubstructMatch`, `countSubstructMatches`, and `getSubstructMatches`.\n\n| Field | Type | Default | Meaning |\n|---|---|---|---|\n| `batchSize` | int | `1024` | (target, query) pairs per GPU batch |\n| `workerThreads` | int | `-1` (auto) | GPU runner threads per GPU |\n| `preprocessingThreads` | int | `-1` (auto) | CPU threads for preprocessing |\n| `maxMatches` | int | `0` (unlimited) | Max matches returned per (target, query) pair |\n| `uniquify` | bool | `False` | Drop duplicate matches that differ only in atom enumeration order |\n| `gpuIds` | `list[int] \\| None` | `None` (current device only) | Specific device ordinals to target |\n\nSubstructure search currently does not support chirality-aware matching, enhanced stereochemistry, or other advanced RDKit `SubstructMatchParameters` options.\n\n## Recipes\n\n### Morgan fingerprints + bulk Tanimoto similarity\n\n```python\nimport torch\nfrom rdkit import Chem\nfrom nvmolkit.fingerprints import MorganFingerprintGenerator\nfrom nvmolkit.similarity import crossTanimotoSimilarity\n\nsmiles = [\"CCO\", \"CCN\", \"c1ccccc1\", \"CC(=O)O\", \"CCOCC\"]\nmols = [Chem.MolFromSmiles(smi) for smi in smiles]\n\nfpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)\nfps = fpgen.GetFingerprints(mols)\n\nsim = crossTanimotoSimilarity(fps)\ntorch.cuda.synchronize()\nprint(sim.torch())\n```\n\nInputs are `list[Mol]`. Output of `GetFingerprints` is an `AsyncGpuResult` wrapping an `(n_mols, fpSize / 32)` int32 tensor of packed bits. Pass it straight into `crossTanimotoSimilarity` for an `(n, n)` similarity matrix; pass two fingerprint sets for an `(n, m)` cross-matrix. For sets too large to materialize on the GPU, use `crossTanimotoSimilarityMemoryConstrained` (chunked compute, returns numpy on CPU).\n\n### ETKDG conformer embedding\n\n```python\nfrom rdkit.Chem import AddHs, MolFromSmiles\nfrom rdkit.Chem.rdDistGeom import ETKDGv3\nfrom nvmolkit.embedMolecules import EmbedMolecules\n\nmols = [AddHs(MolFromSmiles(smi)) for smi in [\"C1CCCCC1\", \"C1CCCCC2CCCCC12\", \"COO\"]]\nparams = ETKDGv3()\nparams.useRandomCoords = True\n\nEmbedMolecules(mols, params, confsPerMolecule=10, maxIterations=-1)\n\nfor mol in mols:\n print(mol.GetNumConformers())\n```\n\nInputs are `list[Mol]`, sanitized and with hydrogens added (`AddHs`). Conformers are added in-place. `params.useRandomCoords` must be `True` - nvMolKit's ETKDG only supports random-coord initialization. A handful of niche `EmbedParameters` options are not supported (bounds matrices, custom CPCI, coord maps, separate-fragment embedding); the Features section of the docs site lists the full restrictions.\n\n### MMFF94 minimization of a batch of conformers\n\n```python\nfrom rdkit.Chem import AddHs, MolFromSmiles\nfrom rdkit.Chem.rdDistGeom import ETKDGv3\nfrom nvmolkit.embedMolecules import EmbedMolecules\nfrom nvmolkit.mmffOptimization import MMFFOptimizeMoleculesConfs\n\nmols = [AddHs(MolFromSmiles(smi)) for smi in [\"CCO\", \"CCN\", \"c1ccccc1\"]]\nparams = ETKDGv3(); params.useRandomCoords = True\nEmbedMolecules(mols, params, confsPerMolecule=5)\n\nenergies = MMFFOptimizeMoleculesConfs(mols, maxIters=500)\nfor mol, mol_energies in zip(mols, energies):\n print(mol.GetNumConformers(), mol_energies)\n```\n\nInputs are `list[Mol]` with conformers already populated (typically by ETKDG, RDKit's `EmbedMultipleConfs`, or a prior nvMolKit call). Coordinates are updated in place; the return is `list[list[float]]` of optimized energies aligned with the input molecule order and conformer index. UFF is identical in shape: swap in `from nvmolkit.uffOptimization import UFFOptimizeMoleculesConfs`.\n\nIf any input molecule is `None` or lacks MMFF/UFF atom types, the call raises `ValueError`. The exception's `args[1]` is a dict with keys `\"none\"` and `\"no_params\"` listing the offending indices - useful for filtering a noisy input set.\n\n### Conformer RMSD and Butina clustering\n\n```python\nimport torch\nfrom rdkit import Chem\nfrom rdkit.Chem.rdDistGeom import EmbedMultipleConfs\nfrom nvmolkit.clustering import butina\nfrom nvmolkit.conformerRmsd import GetConformerRMSMatrixBatch\n\nmols = [Chem.AddHs(Chem.MolFromSmiles(smi)) for smi in [\"CCCCCC\", \"c1ccccc1\"]]\nfor mol in mols:\n EmbedMultipleConfs(mol, numConfs=10)\n\n# Remove hydrogens after embedding for heavy-atom RMSD.\nheavy_mols = [Chem.RemoveHs(mol) for mol in mols]\n\n# Default RMSD output is RDKit-compatible condensed lower-triangle form.\ncondensed = GetConformerRMSMatrixBatch(heavy_mols)\n\n# Butina expects a square distance matrix, so request square GPU tensors.\nsquare = GetConformerRMSMatrixBatch(heavy_mols, output_format=\"square\")\nclusters = [butina(distance_matrix, cutoff=0.5).torch() for distance_matrix in square]\n\ntorch.cuda.synchronize()\nfor mol_clusters in clusters:\n print(mol_clusters.cpu().tolist())\n```\n\n`GetConformerRMSMatrix(mol)` and `GetConformerRMSMatrixBatch(mols)` default to `output_format=\"condensed\"`, returning `AsyncGpuResult` objects that wrap RDKit-style flat vectors of length `N * (N - 1) // 2`. Use `output_format=\"square\"` when chaining into `butina()` or any other API that expects an `N x N` distance matrix. Both forms live on the GPU; call `.numpy()` on condensed results or synchronize before moving square tensors to the CPU.\n\n### Custom forcefield options + constraints (`BatchedForcefield`)\n\nReach for `MMFFBatchedForcefield` / `UFFBatchedForcefield` instead of the one-shot `MMFFOptimizeMoleculesConfs` / `UFFOptimizeMoleculesConfs` when you need any of:\n\n- Custom `maxIters` / `forceTol` per call\n- Per-molecule `nonBondedThreshold` (MMFF) or `vdwThreshold` (UFF), or per-molecule `ignoreInterfragInteractions`\n- Per-molecule `MMFFMolProperties` objects (e.g. MMFF94s vs MMFF94)\n- Distance, position, angle, or torsion constraints\n- Standalone `compute_energy()` / `compute_gradients()` without minimization\n\n```python\nfrom rdkit.Chem import AddHs, MolFromSmiles\nfrom rdkit.Chem.rdDistGeom import EmbedMultipleConfs\nfrom nvmolkit.batchedForcefield import MMFFBatchedForcefield\n\nmols = [AddHs(MolFromSmiles(smi)) for smi in [\"CCO\", \"CCCCCC\"]]\nfor mol in mols:\n EmbedMultipleConfs(mol, numConfs=5)\n\nff = MMFFBatchedForcefield(\n mols,\n nonBondedThreshold=[100.0, 20.0],\n ignoreInterfragInteractions=True,\n)\n\nff[0].add_position_constraint(0, max_displ=0.1, force_constant=50.0)\nff[1].add_distance_constraint(0, 4, relative=False, min_len=1.8, max_len=2.2, force_constant=25.0)\n\nenergies, converged = ff.minimize(maxIters=500, forceTol=1e-4)\nfor mol, mol_energies, mol_converged in zip(mols, energies, converged):\n print(mol.GetNumConformers(), mol_energies, mol_converged)\n```\n\nAll conformers of each input molecule are minimized in one batch. Constraints attached via `ff[i].add_*_constraint(...)` apply to every conformer of molecule `i`; constraint setters mark the wrapper dirty and the native forcefield rebuilds on the next call. Pass `output=CoordinateOutput.DEVICE` to `.minimize(...)` to keep optimized coordinates on the GPU (`Device3DResult`) instead of writing them back into RDKit conformers. UFF is the same shape: `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`.\n\n## Going deeper\n\n- Full feature list, API reference, and guides: <https://nvidia-bionemo.github.io/nvMolKit/>\n- What changed in each release: <https://nvidia-bionemo.github.io/nvMolKit/changelog.html>\n- Worked examples (Jupyter notebooks): the `examples/` directory in the GitHub repo\n"
}SHA-256: dcfef2f4e62ba051e14c12d2fc51e22175e27c930c36f4b6e9ffe3508782225b