NVIDIA BioNeMo Agent Toolkit
NVIDIA BioNeMo v0.1.0
A decade of NVIDIA life-sciences libraries, NIMs, and open models as portable agent skills: structure prediction, docking, generative chemistry, genomics, and de novo protein design. The plugin helps an agent pick the right tool, prepare inputs, run it, and interpret outputs.
Language: English · Automatically detected from descriptions.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- Apache-2.0
- Package author
- NVIDIA BioNeMo
- Keywords
- bionemo, life-sciences, drug-discovery, protein-folding, molecular-docking, generative-chemistry, genomics, protein-design, cuda-x, nim
Declared capabilities
- Interactive
- Write
Package observed Sep 30, 2026.
Files & skills
File archives
Skill instructions
boltz2-nim5.25 KB
---
name: boltz2-nim
description: >
Use Boltz2 NIM for biomolecular structure prediction and binding affinity. Invoke for Boltz2, protein structures, protein-ligand/DNA/RNA complexes, SMILES or CCD ligands, pIC50/IC50 affinity scoring, mmCIF output, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Boltz2 NIM
Predict biomolecular structures and optional ligand affinity. Use this
`SKILL.md` for first-pass hosted/local usage; load supplemental files only when
needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: purpose, strengths, limitations, and handoffs.
- `references/parameters.md`: prediction, sampling, MSA, template, affinity tuning.
- `references/validation.md`: mmCIF, confidence, affinity, and chemistry checks.
- `references/examples.md`: compact hosted/local payload patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/mit/boltz2/predict`
- Local: `http://localhost:8000/biology/mit/boltz2/predict`
Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker
For local setup answers, copy the preflight below before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not invent a cache
default or drop the `.env` load or `NVIDIA_API_KEY` fallback.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm --name boltz2 --gpus all \
--shm-size=16G \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.6.0
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
First startup downloads about 30 GB of model weights.
## Request Pattern
```python
import os
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/mit/boltz2/predict"
if HOSTED else "http://localhost:8000/biology/mit/boltz2/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"polymers": [{
"id": "A",
"molecule_type": "protein",
"sequence": "MTEYKLVVVGACGVGKSALTIQLIQNHFVDEYDPT",
}],
"recycling_steps": 3,
"sampling_steps": 50,
"diffusion_samples": 1,
"step_scale": 1.638,
"output_format": "mmcif",
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
Payload essentials:
- Protein polymer: `{"molecule_type": "protein", "sequence": "..."}`.
- DNA/RNA polymer: add another polymer with `molecule_type` `"dna"` or `"rna"`.
- Ligand by SMILES: `{"id": "L1", "smiles": "CC(=O)OC1=CC=CC=C1C(=O)O"}`.
- Ligand by CCD: `{"id": "L1", "ccd": "ATP"}`.
- Affinity: set `"predict_affinity": True` on exactly one ligand; report
`affinity_pic50`, `affinity_pred_value`, and `affinity_probability_binary`.
- Precomputed A3M MSA goes under the protein polymer. The A3M record uses
`alignment`, `format`, and `rank`; do not use a stale `data` field.
```python
protein_with_msa = {
"id": "A",
"molecule_type": "protein",
"sequence": "MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT",
"msa": {"msa_search": {"a3m": {
"alignment": ">query\nMTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT",
"format": "a3m",
"rank": 0,
}}},
}
```
## Save And Report Output
```python
for i, structure in enumerate(result["structures"], start=1):
with open(f"structure_{i}.cif", "w", encoding="utf-8") as handle:
handle.write(structure["structure"])
for i, score in enumerate(result.get("confidence_scores", []), start=1):
print(f"structure {i} confidence {score:.4f}")
if "affinities" in result:
for ligand_id, aff in result["affinities"].items():
print(ligand_id, aff["affinity_pic50"][0], aff["affinity_pred_value"][0], aff["affinity_probability_binary"][0])
```
Save every `.cif` artifact. Visualize in PyMOL, ChimeraX, or UCSF Chimera. For
confidence/affinity sanity checks, read `references/validation.md`.
## Limits And Troubleshooting
- Polymers/request: 12. Ligands/request: 20. Chain length: 4096 residues.
- Affinity prediction supports one ligand per request and adds runtime.
- `422`: invalid sequence, invalid CCD/SMILES, malformed MSA, or multiple
affinity ligands.
- Local URL/auth: local path has no hosted auth header; wait on `/v1/health/ready`.
- Local startup: use `--gpus all`, `--shm-size=16G`, and the `/opt/nim/.cache` mount.
Referenced files: 5
complexa-binder-design12.5 KB
---
name: complexa-binder-design
description: >
Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided test-time search (best-of-n, beam search, FK steering, MCTS), select with the internal AF2 reward gate, then INDEPENDENTLY validate each binder by refolding the complex with Boltz2 (default) or OpenFold3 and rank on interface confidence, pLDDT, ipSAE, apo/holo stability, and hotspot contact. Use whenever the user wants de novo binders against a named target, sequence, or PDB, hotspot/epitope-targeted design, Proteina-Complexa / Complexa, or ranked validated binders from one request. Sibling of protein-binder-design (RFdiffusion + ProteinMPNN); this skill uses Proteina-Complexa.
license: Apache-2.0
compatibility: "python>=3.10; numpy>=1.24; gemmi (target prep + Boltz2 templates); pyyaml (target registration)"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Complexa Binder Design (workflow)
From one request — "design binders for `<target>`" — to ranked, **independently
validated** binders. Each returned binder is a **co-designed sequence + predicted
binder–target complex**, gated by interface confidence, by whether the binder
actually contacts the target hotspots, and by **apo/holo stability**.
Generation uses **Proteina-Complexa** (co-designs binder sequence + full-atom
structure together — no inverse-folding step — with reward-guided test-time search).
Validation uses a **different** model family (Boltz2 / OpenFold3), so the headline
confidence is an independent check, not the generator grading its own homework.
> **Upstream model + code (you provide these):**
> - Project page: <https://research.nvidia.com/labs/genair/proteina-complexa/>
> - Code: <https://github.com/NVIDIA-Digital-Bio/Proteina-Complexa> (the `complexa` CLI)
> - Weights (NGC): `nvidia/clara/proteina_complexa`
> - Paper: Didi et al., *Scaling Atomistic Protein Binder Design…*, ICLR 2026.
> **First time on a host? → `references/setup.md`** — full standalone setup with **no
> NIM**: install Proteina-Complexa + download weights, Python deps (`numpy gemmi
> pyyaml`), AF2 **configure-vs-bypass**, optional analyze tools (`foldseek`/`sc`/`dssp`),
> the Boltz2/OF3 validation endpoint, and every env var. Then run
> `bash scripts/check_setup.sh` for a one-shot readiness checklist.
```
Stage 1: Resolve target + hotspots → target.pdb + hotspots.json (no GPU)
┌──────────────────────────────────────────────────────────────────────┐
│ repeat until ≥ N validated passers (or a stop cap): │
Stage 2 │ Generate (complexa design) → complex .pdb + AF2-reward-gated designs │
Stage 3 │ Validate (Boltz2 default; OF3 optional) → holo+apo + ipTM/ipSAE/ │
│ pLDDT + apo↔holo RMSD + hotspot contact → passers │
└──────────────────────────────────────────────────────────────────────┘
Stage 4: Report → REPORT_<target>_<run>.md (GO/NO-GO)
```
## Composed pieces (read on demand — do not inline)
| Step | Tool | Owns |
|---|---|---|
| Target + hotspots | vendored `science-skills` (UniProt, AFDB) + `scripts/` | structure resolution, evidence-based hotspots, ≤500 crop, preflight |
| Generation | **Proteina-Complexa** `complexa` CLI | co-designed binder seq+structure, AF2-reward gate → `references/complexa-cli.md` |
| Validate / score | `boltz2-nim` (default) or `openfold3-nim` | independent holo+apo refold, ipTM / pLDDT / PAE |
| MSA (target) | `msa-search-nim` or `scripts/fetch_target_msa_colabfold.py` | target A3M for higher-confidence refolds |
If you run inside the Proteina-Complexa repo, its bundled `.claude/skills`
(`complexa-setup`, `complexa-target`, `complexa-design`) can drive the generation
half; this skill adds the automated Stage 1, the independent validation, GO/NO-GO,
and the manifest.
## Stage 1 — resolve target and hotspots (no GPU)
The user gives a target as a **name**, **sequence**, and/or **structure file**.
Resolve exactly **one design-ready structure**, in priority order: (1) experimental
**PDB** (RCSB), (2) **AFDB** model (UniProt → `vendor/science-skills/.../fetch_structure.py`),
(3) **user-provided** file, (4) **fold de novo** (MSA-Search + OpenFold3/Boltz2).
`scripts/pipeline.py:resolve_target_spec`/`resolve_target` automate (1)–(2) from
free text.
**Hotspots** = the target residues the binder should contact — a compact,
surface-exposed, binder-accessible epitope. Resolve in evidence order
(`scripts/hotspot_strategy.py`, `scripts/pdb_interface.py`):
1. **PDB co-complex interface** (gold standard) — interface residues from a structure
where the target contacts a protein partner.
2. **UniProt functional residues** — `Mutagenesis` + accessible `Active/Binding/Site`,
**filtered to the extracellular/accessible range** (catalytic/cytoplasmic pockets
are the wrong surface for a binder and are dropped).
3. **Literature (Paperclip)** — full-text mining when 1–2 are empty
(`prompts/hotspot_paperclip.md`); structure-confirmed to auto-correct numbering.
4. **Unconditioned** (`[]`) only as a documented last resort.
Then enforce, deterministically:
- **Structure alignment** (`align_hotspots_to_structure`) — drop residues absent from
the coordinate file; read back the real 3-letter identity (catches UniProt↔PDB
numbering offsets — never assume equal indices or chain `A`).
- **Epitope sanity** (`_prune_hotspots`) — one compact patch: drop outliers > 30 Å
from the cluster centroid, cap at 15 residues, prefer ≥ 2.
- **Size budget ≤ 500 residues** (`_crop_target_to_epitope`) — Complexa builds an
O(n²) pair-feature map over the whole complex, so crop large targets to an epitope
window (original numbering preserved).
**Preflight (no GPU):** `python3 scripts/preflight_design.py <name|accession> …`
reports the conditioned length, re-aligned hotspots + source, compactness, the ≤500
budget, and a READY / NEEDS-ATTENTION verdict. Review before spending GPU.
## Stage 2 — generate (Proteina-Complexa, open CLI)
Register the target (hotspots + binder length are target-dict-driven), then **use
`complexa generate` (NOT the full `complexa design`)** for the lean, fast path:
```bash
python scripts/complexa_design.py run --task-name <name> --run-name <run> \
--algorithm best-of-n --num-samples <N> --seed 0 --out <run-dir>
```
`complexa_design.py run` defaults to the **`generate`** verb. With **`best-of-n` + the
AF2 reward** (AF2 params configured via `setup_af2_params.sh` + `AF2_DIR`), the search
**AF2-selects the best candidates during generation** and writes co-designed
**sequence + structure** PDBs to `inference/` — **use the sequence directly, do not
MPNN-redesign it**. Search algorithms: `best-of-n` (default) · `beam-search` ·
`fk-steering` · `mcts`. Overrides + outputs: `references/complexa-cli.md`.
> **Do NOT run the full `complexa design` for this workflow.** Its `evaluate` stage
> **re-folds every design with AF2/RF3/ESMFold (redundant** — best-of-n already
> AF2-selected during search**)** and its `analyze` stage needs `foldseek`/`sc`
> (usually not installed). It is much slower and adds a failure mode. The lean
> `generate` → independent **Boltz2** validation (Stage 3) is the intended path.
>
> No AF2 params? use `--af2-bypass` (`single-pass` + drop the AF2 reward); selection
> then falls entirely to the independent Boltz2 gate (Stage 3). Low-complexity
> (poly-X) sequences are dropped before spending Boltz2.
## Stage 3 — validate (independent refold) + gate
**Validate a capped shortlist, not the whole pool.** Best-of-n produces many
candidates; fold only ~**2× the requested N** (the top ones by the generation/AF2
reward) — validating the entire pool wastes GPU/time and (on hosted Boltz2) trips rate
limits. Point at a local Boltz2 NIM via `$BOLTZ2_URL` (`--endpoint local`) when available.
Per binder run **two** predictions with one refolder (Boltz2 default): **holo**
(binder + target; target MSA, binder single-sequence, `write_full_pae`) and **apo**
(binder alone). One command does it: **`scripts/boltz2_refold.py`** makes the holo
calls (with retry/backoff for rate limits) and chains **`scripts/validate_binders.py`**,
which runs apo + computes the metrics + applies the gate + ranks. Per-chain
conditioning + metric definitions: `references/validation.md`.
**Gate (defaults — every gate must hold):** ipTM ≥ 0.65, complex pLDDT ≥ 0.70, binder
pLDDT ≥ 0.70, apo binder pLDDT ≥ 0.70, **ipSAE_min ≥ 0.45**, apo↔holo binder RMSD
≤ 2.5 Å, ≥ 20% of conditioned hotspots contacted (CB–CB < 13 Å). Record **every**
design (pass *and* fail) with a `failure_reason`. Rank protein binders by interface
confidence (ipTM/ipSAE) + pLDDT + stability — **not** Boltz2 `affinity_pic50`
(ligand-only).
## Bounded-budget loop + report
The deliverable is **the top-N binders ranked by interface confidence** (default 10).
Aim for N that pass the full gate, but **bound the cost**: run **at most 2 generation
rounds**, then **deliver the top-N by score (ipTM, then ipSAE_min) even if fewer than N
clear the strict gate** — keep each design's `pass`/`failure_reason` flag so quality is
still visible. Do **not** keep generating just to chase N strict passes (that is the
single biggest time sink). Stop on N-passed / 2 rounds / budget / a zero-passer round.
One run dir per campaign; `manifest.json` records target, Complexa run config + seeds,
per-design lineage/scores/artifacts, gate status. The report states GO/NO-GO,
requested-vs-achieved N (passed and delivered), ranked binders, and which stop
condition fired. Layout, loop, and report sections: `references/pipeline.md`.
## Configuration
- `COMPLEXA_REPO` — path to your local Proteina-Complexa checkout (the `complexa`
CLI runs there). Checkpoints via the pipeline YAML or `++ckpt_path=…`. Reward
weights (`AF2_DIR`, `RF3_CKPT_PATH`/`RF3_EXEC_PATH`) via the repo's `.env`.
- Boltz2 / OpenFold3 endpoints + auth: hosted (`https://health.api.nvidia.com/v1/…`
+ `NVIDIA_API_KEY`) or local (`http://localhost:8000/…`, no auth). The validator
takes `--endpoint hosted|local`; the key is read from `NVIDIA_API_KEY`/`NGC_API_KEY`
(or `--env-file`). Never hardcode hosts/keys.
- `COMPLEXA_OUTPUTS` — run-output root (default `./outputs`).
## Scripts & assets
- `references/setup.md` + `scripts/check_setup.sh` — standalone (no-NIM) install guide
and a one-shot environment readiness check.
- `scripts/pipeline.py` — orchestrator (Stage-1 resolution + open-CLI generation +
AF2 gate + scoring); `score_existing` and `full` modes.
- `scripts/preflight_design.py` — no-GPU target/hotspot/size planner.
- `scripts/hotspot_strategy.py`, `scripts/pdb_interface.py` — evidence-based hotspots.
- `scripts/complexa_design.py` — thin `complexa design` driver + output extraction.
- `scripts/setup_af2_params.sh` — download AF2-Multimer params (public, no auth) +
create the `params/` layout, for reward-guided search (`best-of-n`, etc.).
- `scripts/boltz2_refold.py` — Stage-3 **holo** Boltz2 refolds (retry/backoff + throttle)
→ `validation/raw/`, then chains `validate_binders.py`.
- `scripts/validate_binders.py` — apo Boltz2 + scoring, ipSAE, apo↔holo RMSD,
hotspot contact, gating, ranking → `ranked_binders.json`. Needs the Dunbrack **ipSAE**
script: `bash scripts/fetch_ipsae.sh` (MIT; fetched, not bundled — see `vendor/ipsae/`).
- `scripts/fetch_target_msa_colabfold.py`, `scripts/pdb_to_boltz_template_cif.py` —
target MSA / structural-template helpers for validation.
- `vendor/science-skills/` — DeepMind UniProt + AFDB tooling (Apache-2.0) for Stage 1.
- `prompts/hotspot_paperclip.md` — literature-mining hotspot fallback.
- `assets/targets.json` — example registered targets (use `complexa target add` for your own).
## Responsible use
De novo binder design is dual-use. Decline requests aimed at enhancing pathogen
fitness, toxin potency, or bioweapon function; keep designs to legitimate research
and therapeutic intent.
## See also
- `protein-binder-design` — same goal via RFdiffusion + ProteinMPNN (BioNeMo NIMs).
- Proteina-Complexa docs: `README.md`, `docs/INFERENCE.md`, `docs/CONFIGURATION_GUIDE.md`,
`docs/EVALUATION_METRICS.md`, and its bundled `.claude/skills/` in the repo above.
Referenced files: 32
complexa-design16.2 KB
---
name: complexa-design
description: >
End-to-end Proteina-Complexa design pipeline driver. Reach for this skill whenever
the user wants to "design a binder", "design binders for X", "run complexa
design", "de novo binder", "PDL1 binder", "TrkA binder", "design proteins for
target", "protein binder design", "ligand binder", "design a small-molecule
binder", "ATP-binding protein", "AME motif scaffolding", "scaffold a motif
near a ligand", "motif + ligand design", "enzyme scaffolding", "flow matching
protein design", "beam-search binder", "FK steering", "MCTS protein design",
"refold with AF2", "refold with RF3", or wants success rates, interface pAE,
scRMSD, or
FoldSeek diversity from a single command. This is the scientific anchor of
the skill set: it drives `complexa design <pipeline>` from target picking to
manifest emission and tells the user how many designs passed.
compatibility: "complexa CLI installed (pip install -e .); .env populated; 1x CUDA GPU >=40GB VRAM (A100/H100/L40S); 24 CPUs; ~50GB disk"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Complexa Design Skill
Drive the full four-stage `complexa design` pipeline: generate (flow matching +
search) -> filter (top-N by reward) -> evaluate (refold with AF2/RF3) ->
analyze (success rate, FoldSeek/MMseqs diversity). Pick the right pipeline
config for the design intent, validate the run upfront so the user does not
discover a missing ckpt mid-folding, run it, and emit a replayable manifest +
per-design success CSV.
## What this skill enables
- Protein binder design for protein targets (AF2 reward + ColabDesign refold).
- Ligand binder design for small-molecule targets (RF3 reward + RF3 refold).
- AME motif scaffolding with ligand context (motif + ligand features, RF3).
- Search-based optimization: single-pass, best-of-n, beam-search, fk-steering, mcts.
- Refold backends: ColabDesign (AF2), RF3, Boltz2, ESMFold (fast iteration).
- Pass-rate + diversity analysis with per-`result_type` thresholds.
## Step 1: Pre-flight
Always run the shared preflight before launching a design — generation needs the
GPU and the right checkpoint, evaluation needs AF2/RF3 weights and tool
binaries. Bail early if the host cannot run the chosen pipeline.
```bash
bash .claude/skills/_shared/scripts/preflight.sh
```
Read `./complexa_setup/preflight.json` and bail if any of these are missing for
the chosen pipeline:
- `gpu.available: false` -> all pipelines fail.
- `gpu.vram_gb < 40` -> generation OOMs at default `batch_size: 16`; lower to 8.
- `ckpts.complexa[.ckpt]` -> required for protein binder.
- `ckpts.complexa_ligand[.ckpt]` -> required for ligand binder.
- `ckpts.complexa_ame[.ckpt]` -> required for AME.
- `env.AF2_DIR` missing -> protein binder default eval (`colabdesign`) fails.
- `env.RF3_CKPT_PATH` or `env.RF3_EXEC_PATH` missing -> ligand binder / AME default eval (`rf3_latest`) fails.
If a ckpt is missing, point at `complexa-setup` and have the user run
`complexa download --complexa-<variant>` first.
## Step 2: Pick the pipeline
Complexa has **one default pipeline (protein binder) and two extensions**
(ligand binder, AME / enzyme). The pipeline is selected entirely
by the `configs/search_*_pipeline.yaml` you pass to `complexa design` — each
YAML pins its own model checkpoint, autoencoder, targets dict, default reward,
and default refold backend. Switching pipelines is just "swap the config path
and the target name comes from a different dict".
### Default — protein binder
```bash
complexa design configs/search_binder_local_pipeline.yaml \
++run_name=pdl1_v1 ++generation.task_name=02_PDL1
```
Use this when the user says "design a binder for X", "PDL1 binder",
"de novo binder", "design proteins for a target", etc. — i.e. the target is
a protein surface. The config pins `complexa.ckpt` + `complexa_ae.ckpt`, reads
targets from `configs/targets/targets_dict.yaml`, rewards with AF2
(`af2folding`), inverse-folds with SolubleMPNN, and evaluates with
ColabDesign / AF2. **If the user did not specify, this is what they want.**
### Extension A — ligand binder (small-molecule pocket)
```bash
complexa design configs/search_ligand_binder_local_pipeline.yaml \
++run_name=v11_v1 ++generation.task_name=39_7V11_LIGAND \
++metric.binder_folding_method=rf3_latest
```
Switch to this when the user says "ligand binder", "small-molecule pocket",
"SMILES target", "ATP-binding protein", or names a target ending in
`_LIGAND` / from the FAD / SAM / OQO / 7V11 / etc. families. The config
**also activates LoRA** (`r=32`, `lora_alpha=64`) which is required for the
released ligand checkpoint — leave the `lora:` block alone.
### Extension B — AME / motif + ligand (enzyme scaffolding)
```bash
complexa design configs/search_ame_local_pipeline.yaml \
++run_name=ame_chm ++generation.task_name=M0096_1chm
```
Switch to this when the user says "scaffold a motif near a ligand",
"active-site design", "enzyme scaffolding", "AME", or names a target like
`M0024_1nzy`, `M0096_1chm` (the `M####_<pdb>` AME task naming). The config sets
`env_vars.USE_V2_COMPLEXA_ARCH=True` (the CLI runner injects it into the
subprocess) and uses both `MotifFeatures` and `LigandFeatures`. Default search
is `single-pass`; switch to `best-of-n` only if you also enable the
`CompositeRewardModel` (commented out in `ame_generate.yaml`).
### Pipeline cheat sheet — what changes when you switch
| Knob | Protein binder (default) | Ligand binder | AME (enzyme) |
|---|---|---|---|
| **Pipeline YAML** | `configs/search_binder_local_pipeline.yaml` | `configs/search_ligand_binder_local_pipeline.yaml` | `configs/search_ame_local_pipeline.yaml` |
| **Model ckpt** | `complexa.ckpt` | `complexa_ligand.ckpt` | `complexa_ame.ckpt` |
| **Autoencoder ckpt** | `complexa_ae.ckpt` | `complexa_ligand_ae.ckpt` | `complexa_ame_ae.ckpt` |
| **Targets dict** | `configs/targets/targets_dict.yaml` | `configs/targets/ligand_targets_dict.yaml` | `configs/design_tasks/ame_dict_v2.yaml` |
| **Task-name pattern** | `<NN>_<NAME>` (e.g. `02_PDL1`, `22_DerF21`) | `<NN>_<PDB>_LIGAND` (e.g. `39_7V11_LIGAND`) | `M####_<pdb>` (e.g. `M0096_1chm`) |
| **`USE_V2_COMPLEXA_ARCH`** | (unset → v1) | (unset → v1) | `"True"` (set in YAML) |
| **LoRA** | (none) | required (`r=32, alpha=64`) | required |
| **Default search algo** | `best-of-n` | `best-of-n` | `single-pass` |
| **Reward model** | AF2 (`af2folding`) | RF3 (`rf3folding`) | `null` (no reward at default) |
| **Inverse folder** | `soluble_mpnn` | `ligand_mpnn` | `ligand_mpnn` |
| **Default refold backend** | `colabdesign` (AF2) | `rf3_latest` | `rf3_latest` |
| **Analysis `result_type`** | `protein_binder` | `ligand_binder` | `motif_ligand_binder` |
| **Required ckpts (`complexa download`)** | `--complexa --all` (AF2 in community) | `--complexa-ligand --all` (RF3 in community) | `--complexa-ame --all` (RF3 in community) |
For the full per-pipeline breakdown (reward weights, success thresholds,
analysis modes), see [reference/pipelines.md](reference/pipelines.md).
## Step 3: Gather parameters
Use AskUserQuestion to fill in the four parameters that vary every run. Default
to sensible production settings if the user has no preference.
- **Target name** — must be a key in the relevant dict (`targets_dict.yaml` for
protein binder, `ligand_targets_dict.yaml` for ligand, `ame_dict_v2.yaml` for
AME). If the user names a target that is not in the dict, hand off to
`complexa-target` to add it first.
- **Run name** — a short identifier appended to the output dir (e.g. `pdl1_v1`).
- **Search algorithm** — default to `beam-search` with `beam_width=8` and
`n_branch=4` for production. Use `single-pass` for a quick smoke test.
- **Evaluation refold backend** — protein binder defaults to `colabdesign`
(AF2); ligand/AME default to `rf3_latest`. Use `esmfold` for fast iteration
(worse but seconds per sample).
## Step 4: Validate
Validate before running. This is cheap (seconds) and catches missing ckpts,
missing env vars, unknown override keys, and missing target entries — all of
which would otherwise abort the pipeline mid-evaluation after hours of
generation.
```bash
complexa validate design configs/search_binder_local_pipeline.yaml \
++generation.task_name=02_PDL1 \
++metric.binder_folding_method=colabdesign
```
The validator returns non-zero on failure and prints a pass/fail report.
Re-run the command with the suggested overrides until it returns clean.
## Step 5: Run the pipeline
`complexa design` is the right tool for the full 4-stage run — it orchestrates
`generate → filter → evaluate → analyze` as sequential subprocesses with a
shared run name, log dir, and multi-GPU split (see `run_design_pipeline` in
`src/proteinfoundation/cli/cli_runner.py`). Re-implementing that manually
loses the per-stage log routing and progress prints.
Use `++` (forced) Hydra overrides; they apply to all stages. The minimal
production protein-binder invocation:
```bash
complexa design configs/search_binder_local_pipeline.yaml \
++run_name=pdl1_v1 \
++generation.task_name=02_PDL1 \
++generation.search.algorithm=beam-search \
++generation.search.beam_search.beam_width=8 \
++metric.binder_folding_method=colabdesign
```
For ligand binder / AME, swap the pipeline YAML and target name per Step 2's
cheat sheet — every other override above is pipeline-agnostic and can be
reused as-is.
Add `--verbose` to stream logs to the terminal instead of `./logs/`. The skill
does not poll progress — the user re-invokes if they want a status; point them
at `complexa status` and `./logs/design_pipeline_*/`.
### Direct module invocation (debug fallback)
To debug a single stage without pipeline orchestration, invoke the underlying
Hydra module directly. `complexa generate CONFIG` is just a logged subprocess
wrapper around this:
```bash
python -m proteinfoundation.generate \
--config-path "$(realpath configs)" \
--config-name search_binder_local_pipeline \
++run_name=debug_pdl1 \
++generation.task_name=02_PDL1
```
Same pattern for `proteinfoundation.{filter,evaluate,analyze}`. Use this when
you want to attach `ipdb`, run under `nsys`, or skip the pipeline log dir.
Prefer `complexa generate/filter/evaluate/analyze` for normal one-shot runs
(you get logging and parallel job splitting for free).
**AME-specific gotcha (when running AME with RF3 refold)**: RF3 will try to
complete missing atoms on the ligand based on its CCD code, which produces
shape errors in RMSD calculations and the wrong structure. Before RF3 sees the
PDB, rename the ligand residue to `L:0` so RF3 treats it as a generic ligand
and skips atom completion. If your AME generation outputs already encode the
ligand this way (the canonical Complexa pipeline does), no extra step is
needed; otherwise patch each PDB with `atomworks.io`:
```python
from atomworks.io import load_any, to_pdb_file
atom_array = load_any("my_design.pdb")[0]
ligand_mask = atom_array.chain_id == "A"
atom_array.res_name[ligand_mask] = "L:0"
to_pdb_file(atom_array, "my_design_rf3_ready.pdb")
```
Same applies if you're piping AME outputs into the `complexa-evaluate-pdbs`
skill with an RF3 backend. Skip the rename for AF2/colabdesign refold or for
non-AME pipelines.
Wall-clock at default (`nsteps=400`, `beam_width=8`, `batch_size=16`, 100
designs, colabdesign eval) is ~30–120 minutes on a single A100/H100.
## Step 6: Collect results
Outputs land in two directories. Surface both:
```bash
ls ./inference/${CONFIG_STEM}_${TASK}_*${RUN_NAME}/ # generated PDBs + filter
ls ./evaluation_results/${RUN_NAME}/ # per-design CSV + analysis
```
Read the combined results CSV and summarize:
```bash
ls ./evaluation_results/*/binder_results_*_combined.csv
ls ./evaluation_results/*/motif_binder_results_*_combined.csv # AME
ls ./evaluation_results/*/res_filter_*_pass_*.csv # success rate
ls ./evaluation_results/*/res_div_foldseek_*.csv # FoldSeek diversity
```
Pull the success rate from `res_filter_binder_pass_*.csv`, the per-design
metrics (interface pAE, pLDDT, scRMSD) from the combined CSV, and FoldSeek
TM-score diversity from `res_div_foldseek_*.csv`. Report top-N designs by
i_pAE (protein binder) or min_ipAE (ligand binder).
## Step 7: Emit manifest
Drop a JSON manifest beside the results so the run is replayable. The shared
helper captures the command, config, git SHA, and pointers to the result CSVs.
```bash
python3 .claude/skills/_shared/scripts/write_manifest.py \
--output-dir ./evaluation_results/${RUN_NAME} \
--command "complexa design configs/search_binder_local_pipeline.yaml ++run_name=${RUN_NAME} ++generation.task_name=${TASK}" \
--skill complexa-design \
--out ./run_manifest.json
```
Surface the manifest path and the result CSV to the user.
## Most-common overrides
The 10 overrides that cover ~90% of runs. Full reference (every key, type,
default) is in [reference/overrides.md](reference/overrides.md).
| Override | Default | What it controls |
|----------|---------|------------------|
| `++generation.task_name=<name>` | (per config) | Which target / AME task to design for |
| `++run_name=<str>` | (config stem) | Output dir suffix and CSV tag |
| `++generation.search.algorithm=beam-search` | `best-of-n` (binder/ligand), `single-pass` (AME) | Search strategy |
| `++generation.search.beam_search.beam_width=8` | `4` | Beam-search width (more = better designs, slower) |
| `++generation.args.nsteps=200` | `400` | Diffusion steps (fewer = faster, lower quality) |
| `++generation.dataloader.batch_size=8` | `16` (binder/ligand/AME) | Drop to 8 on a 40GB GPU |
| `++generation.filter.filter_samples_limit=500` | `1000` | Top-N samples to keep after filtering |
| `++metric.binder_folding_method=esmfold` | `colabdesign` (binder), `rf3_latest` (ligand/AME) | Evaluation refold backend |
| `++metric.num_redesign_seqs=8` | `2` | ProteinMPNN/LigandMPNN/SolubleMPNN sequences per design |
| `++aggregation.success_thresholds.i_pAE.threshold=10.0` | `7.0` (protein binder) | Loosen / tighten success criteria |
## Hardware requirements
| Resource | Minimum | Recommended |
|----------|---------|-------------|
| GPU | 1x CUDA GPU, 40 GB VRAM | A100 / H100 / L40S, 80 GB VRAM |
| CPUs | 16 | 24 (the `ncpus_` default in every pipeline config) |
| Disk | 50 GB at `./inference/` + `./evaluation_results/` | 200 GB for sweep runs |
| RAM | 32 GB | 64 GB+ |
Typical wall-clock for 100 designs, `beam_width=8`, default `nsteps=400`:
- Protein binder + colabdesign refold: ~60–120 min on 1x A100/H100.
- Ligand binder + RF3 refold: ~90–180 min (RF3 dominates).
- AME + RF3 refold: ~120–240 min.
- Any pipeline + ESMFold refold: ~30–60 min (fast iteration).
Bumping `gen_njobs=2` and `eval_njobs=2` halves wall-clock on a 2-GPU host. See
`.claude/skills/_shared/reference/hardware.md` for per-pipeline VRAM tables.
## Troubleshooting (common cases)
| Symptom | Cause | Fix |
|---------|-------|-----|
| `CUDA out of memory` in generate | `batch_size: 16` too big on 40GB GPU | `++generation.dataloader.batch_size=8` |
| `CUDA out of memory` in evaluate | AF2 / RF3 batched too aggressively | `++eval_njobs=1` and `++metric.num_redesign_seqs=2` |
| `InterpolationKeyError: AF2_DIR` | colabdesign eval but `.env` does not set `AF2_DIR` | Set `AF2_DIR` in `.env` or `++metric.binder_folding_method=esmfold` |
| `InterpolationKeyError: RF3_CKPT_PATH` | RF3 eval but RF3 not installed | `complexa download --all` or switch eval backend |
| `KeyError: 'task_name' not in target_dict_cfg` | Target not in `targets_dict.yaml` / `ligand_targets_dict.yaml` / `ame_dict_v2.yaml` | Use `complexa-target` skill to add it |
| 0 designs pass success thresholds | Defaults too strict for this target | Loosen via `++aggregation.success_thresholds.*` |
For the full list (chain-ID mismatches, hotspot residues, ligand residue
renaming for RF3, missing inverse-folding models, etc.) see
[reference/troubleshooting.md](reference/troubleshooting.md).
---
For per-pipeline details (which model, reward, inverse folding, evaluation
backend, `result_type`, LoRA settings, `USE_V2_COMPLEXA_ARCH` toggle), see
[reference/pipelines.md](reference/pipelines.md).
For the full override reference (every `generation.*`, `metric.*`,
`aggregation.*` key with type, default, example, and effect), see
[reference/overrides.md](reference/overrides.md).
For all troubleshooting cases (cause + fix + source), see
[reference/troubleshooting.md](reference/troubleshooting.md).
Referenced files: 3
complexa-evaluate-pdbs13.2 KB
---
name: complexa-evaluate-pdbs
description: >
Standalone evaluation of an existing PDB directory with Proteina-Complexa.
Use this skill whenever the user wants to "evaluate PDB files", "re-fold these
designs", "compute interface pAE", "compute i_pLDDT for a folder",
"run AF2 / RF3 / ESMFold on my designs", "score binder candidates",
"designability of this folder", "scRMSD for designs", "motif RMSD for these
PDBs", "complexa analysis", "complexa evaluate from a PDB directory",
"evaluate from pdb dir", or score third-party outputs (BindCraft, AlphaProteo,
RFdiffusion, hand-curated decoys). It picks the correct `evaluate_*.yaml`
config, wires `++dataset.pdb_dir` and the folding backend, runs
`complexa analysis` (the evaluate → analyze chain), parses the result CSV,
reports pass-rates against the right `result_type` thresholds, and emits a
replayable `eval_manifest.json`. Reach for this skill before hand-rolling
refolding scripts.
compatibility: "complexa CLI installed (pip install -e .); CUDA GPU; AF2_DIR (colabdesign) or RF3_CKPT_PATH+RF3_EXEC_PATH (rf3_latest); ESMFold weights for monomer paths"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Complexa Evaluate-PDBs Skill
Score a directory of pre-existing PDB files against the same metrics Proteina-Complexa uses internally. Wraps `complexa analysis <evaluate_config> ++sample_storage_path=<dir>`: the CLI runs the `evaluate` step (refold + interface metrics + monomer metrics) and then the `analyze` step (success thresholds, diversity, pass-rate CSVs). Do **not** run `complexa generate` here — the inputs already exist.
## What this skill enables
- Re-fold a directory of designed PDBs with AF2 (`colabdesign`), RF3 (`rf3_latest`), ESMFold (`esmfold`), or Boltz2 (`boltz2_default`).
- Compute binder interface metrics: `i_pAE`, `min_ipAE`, `i_pTM`, `pLDDT`, binder/complex scRMSD.
- Compute monomer **designability** (ProteinMPNN-redesigned scRMSD) and **codesignability** (original sequence refold scRMSD).
- For motif inputs: motif RMSD (CA + all-atom), motif-region designability/codesignability, sequence recovery.
- Aggregate into per-PDB CSVs plus pass-rate summaries using the default thresholds for the `result_type`.
## Step 1: Pre-flight
Always check GPU / disk / tool binaries before launching a refold job. RF3 and ColabDesign-AF2 are large.
```bash
bash .claude/skills/_shared/scripts/preflight.sh
```
Surface from `preflight.json`:
- `gpu.available` and `gpu.vram_gb` — colabdesign/RF3 need ≥40 GB; ESMFold tolerates ≥24 GB.
- `env.missing_required` — must include the keys for the chosen folding backend:
- `colabdesign` → `AF2_DIR`
- `rf3_latest` → `RF3_CKPT_PATH`, `RF3_EXEC_PATH`
- `esmfold` → ESMFold weights resolvable
- `tools.{foldseek,mmseqs}` — required by `aggregation.compute_diversity` / `compute_mmseqs_diversity` (both default `true`).
If any required key is missing, route the user to the `complexa-setup` skill.
## Step 2: Identify the design type
Like `complexa design`, evaluation has one default flow (protein binder) and two extensions (ligand binder, AME). The evaluate config you pass to `complexa analysis` decides everything else (which metrics, which refolder defaults, which thresholds the analyze step applies).
### Default — protein binder
```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml \
++sample_storage_path=/abs/path/to/pdbs \
++dataset.task_name=02_PDL1 \
++result_type=protein_binder \
++metric.binder_folding_method=colabdesign \
++metric.inverse_folding_model=soluble_mpnn \
++run_name=eval_pdl1_af2
```
Use this when the user's PDBs are protein-binder designs (multi-chain, binder is the last chain) or third-party outputs from BindCraft / AlphaProteo / RFdiffusion. Pulls thresholds for `protein_binder` (`i_pAE * 31 ≤ 7.0`, `pLDDT ≥ 0.9`, `scRMSD_ca < 1.5 Å`).
### Extensions — pick the matching config
| Design type | Use the protein-binder default? | Evaluate config | Analyze config | `result_type` | Default backend |
|---|---|---|---|---|---|
| Protein binder | **Yes (default)** | `configs/evaluate_from_pdb_dir.yaml` | `configs/analyze.yaml` | `protein_binder` | `colabdesign` (AF2) |
| Ligand binder (binder + small-molecule) | Same evaluate config, swap 3 overrides | `configs/evaluate_from_pdb_dir.yaml` | `configs/analyze.yaml` | `ligand_binder` | `rf3_latest` |
| AME / motif + ligand (enzyme outputs) | No — needs motif-aware config | `configs/evaluate_ame_from_pdb_dir.yaml` | `configs/analyze_motif_binder.yaml` | `motif_ligand_binder` | `rf3_latest` |
| Motif protein binder (standalone) | No — no `_from_pdb_dir` variant | `configs/evaluate_motif_binder.yaml` + `++input_mode=pdb_dir` | `configs/analyze_motif_binder.yaml` | `motif_protein_binder` | `colabdesign` / `rf3_latest` |
**Extending to ligand binder** (same evaluate config as default, three override swaps):
```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml \
++sample_storage_path=/abs/path/to/pdbs \
++dataset.task_name=39_7V11_LIGAND \
++result_type=ligand_binder \
++metric.binder_folding_method=rf3_latest \
++metric.inverse_folding_model=ligand_mpnn \
++run_name=eval_v11_rf3
```
**Extending to AME** (different config; ligand auto-completion gotcha — see Step 4):
```bash
complexa analysis configs/evaluate_ame_from_pdb_dir.yaml \
++sample_storage_path=/abs/path/to/pdbs \
++dataset.task_name=M0096_1chm \
++run_name=eval_ame_chm
```
See `reference/eval_configs.md` for the full matrix (every `result_type`, every threshold default, every supported folding backend, motif-protein-binder variant).
## Step 3: Gather inputs (AskUserQuestion)
Ask in one batched `AskUserQuestion`:
1. **`pdb_dir`** — absolute path to the directory of PDBs to evaluate.
2. **Design type** — protein binder / ligand binder / AME (motif + ligand).
3. **Folding backend** — `colabdesign` (AF2, protein binders), `rf3_latest` (ligand / AME), `boltz2_default` (alternative).
4. **Target / task name** — must match a key in `configs/targets/targets_dict.yaml`, `configs/targets/ligand_targets_dict.yaml`, or `configs/design_tasks/ame_dict_v2.yaml`. Required for binder + AME evaluation (needed to identify the target reference and, for AME, the motif contigs).
5. **AME-only**: confirm ligand residue name is already renamed to `L:0` in every PDB (see Troubleshooting). If not, do that rename first.
## Step 4: Run evaluate → analyze
Prefer `complexa analysis` (the evaluate→analyze chain) — it reuses the same config for both steps and writes a single log dir.
```bash
# Protein binder PDB dir, AF2 refold
complexa analysis configs/evaluate_from_pdb_dir.yaml \
++sample_storage_path=/abs/path/to/pdbs \
++dataset.task_name=02_PDL1 \
++metric.binder_folding_method=colabdesign \
++metric.inverse_folding_model=soluble_mpnn \
++result_type=protein_binder \
++run_name=eval_pdl1_af2
```
For ligand binders flip `binder_folding_method=rf3_latest`, `inverse_folding_model=ligand_mpnn`, `result_type=ligand_binder`. For AME use `configs/evaluate_ame_from_pdb_dir.yaml` — see `reference/eval_configs.md` for full worked examples.
If you need to inspect output between stages, run them separately. The configs above are shared between `evaluate` and `analyze`:
```bash
complexa evaluate configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=/abs/path/to/pdbs ++run_name=eval_pdl1_af2
complexa analyze configs/evaluate_from_pdb_dir.yaml ++run_name=eval_pdl1_af2
```
Dry-run first if the user is unsure (no GPU work happens; the planned file walk + invocation prints):
```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=/abs/path/to/pdbs ++dryrun=true
```
### Direct module invocation (debug fallback)
`complexa evaluate` / `analyze` are subprocess wrappers around the Hydra
modules with logging + parallel job splitting bolted on. To attach a debugger
or run under a profiler, invoke the module directly:
```bash
python -m proteinfoundation.evaluate \
--config-path "$(realpath configs)" \
--config-name evaluate_from_pdb_dir \
++sample_storage_path=/abs/path/to/pdbs \
++dataset.task_name=02_PDL1 \
++metric.binder_folding_method=colabdesign \
++run_name=eval_debug
```
For normal one-shot runs prefer `complexa analysis` — you get the shared log
dir and a single replayable invocation, instead of having to thread the same
overrides through two `python -m` calls.
## Step 5: Parse results
Output lands under `./evaluation_results/${run_name}/`:
- Per-PDB metrics CSV — `*_results_*.csv` (one row per input PDB × `sequence_types`).
- Pass-rate summaries — written by the `analyze` step (e.g. `res_designability.csv`, `res_filter_ligand_pass_*.csv`, `success_criteria_*.json`).
- Diversity output — FoldSeek/MMseqs2 cluster files when `aggregation.compute_diversity=true` (default).
Summarize to the user:
- Per-PDB row count and number of successful designs vs total.
- Default-threshold pass rate by `result_type` (e.g. for `protein_binder`: `i_pAE*31 <= 7.0 AND pLDDT >= 0.9 AND scRMSD_ca < 1.5`).
- Top 5 designs by primary metric (`i_pAE` for protein, `min_ipAE` for ligand, `motif_rmsd_pred_all` for motif binders).
## Step 6: Emit manifest
Capture the resolved invocation + outputs for replay.
```bash
python3 .claude/skills/_shared/scripts/write_manifest.py \
--output-dir ./evaluation_results/${run_name} \
--command "complexa analysis configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=<dir> ++dataset.task_name=<task> ++metric.binder_folding_method=<backend> ++run_name=<run>" \
--skill complexa-evaluate-pdbs \
--out ./eval_manifest.json
```
The manifest pins: resolved config, git SHA, ckpt SHA-256s, the result CSV paths, and the user-stated `result_type` for replay.
## Most common overrides
| Override | Effect |
|------------------------------------------------|-----------------------------------------------------------------------|
| `++sample_storage_path=<dir>` | The directory of PDBs to evaluate (required). |
| `++dataset.task_name=<name>` | Target / AME task name (binders, AME). Resolves target PDB + contigs. |
| `++metric.binder_folding_method=<backend>` | `colabdesign` / `rf3_latest` / `boltz2_default`. |
| `++metric.inverse_folding_model=<model>` | `protein_mpnn` / `soluble_mpnn` / `ligand_mpnn`. |
| `++metric.sequence_types=[self,mpnn,mpnn_fixed]` | Which sequence flavors to refold. |
| `++metric.num_redesign_seqs=N` | ProteinMPNN/LigandMPNN redesign count. |
| `++metric.compute_pre_refolding_metrics=true` | Add bioinformatics/TMOL/HBPLUS metrics on the input structures. |
| `++metric.keep_folding_outputs=true` | Save the refolded PDBs (large, but useful for inspection). |
| `++result_type=<type>` | Override default thresholds: `protein_binder` / `ligand_binder` / `motif_protein_binder` / `motif_ligand_binder`. |
| `++aggregation.success_thresholds.<…>` | Tighten or loosen specific thresholds (see `reference/eval_configs.md`). |
| `++eval_njobs=N` | Parallel GPUs for the evaluate step. |
| `++dryrun=true` | Plan without running any folding. |
| `++file_limit=N` | Cap input PDBs (handy for first-pass smoke tests). |
## Hardware
- **GPU**: ≥1 CUDA GPU. AF2 (`colabdesign`) and RF3 (`rf3_latest`) need ≥40 GB VRAM (A100/H100/L40S). ESMFold runs on ≥24 GB. Multi-GPU via `++eval_njobs=N`.
- **CPU/disk**: 24 CPUs default (`ncpus_: 24`). Each refolded PDB + intermediate output is ~1–5 MB; `keep_folding_outputs=true` can balloon to tens of GB for thousands of inputs.
- See `_shared/reference/hardware.md` for per-backend wall-clock and VRAM tables.
## Troubleshooting
- **`Error: Config file not found`** — paths are relative to the repo root; `cd` to the repo before invoking `complexa analysis`.
- **`compute_motif_binder_metrics=True` but `result_type=protein_binder`** — `result_type` and the underlying `compute_*_metrics` must agree. Use `configs/evaluate_ame_from_pdb_dir.yaml` for AME inputs rather than mutating `evaluate_from_pdb_dir.yaml`.
- **RF3 shape errors on AME PDBs** — RF3 tries to auto-complete the ligand atoms from CCD. Rename the ligand residue to `L:0` in every input PDB before evaluation; see the snippet in `README.md` (`atom_array.res_name[ligand_mask] = "L:0"`).
- **Diversity step fails (`foldseek not found`)** — `FOLDSEEK_EXEC` not on `PATH`. Either fix `.env` (preferred) or disable: `++aggregation.compute_diversity=false ++aggregation.compute_mmseqs_diversity=false`.
- **All pass-rates are 0%** — check the `binder_folding_method` matches the target type (RF3 for ligand, AF2 for protein) and that `++dataset.task_name` resolves to the correct reference PDB (`complexa target show <name>` to verify).
## Reference
Full evaluate/analyze config matrix, every supported `result_type`, per-threshold defaults, and three worked examples (protein binder / ligand binder / AME): see `reference/eval_configs.md`.
Referenced files: 1
complexa-setup13.7 KB
---
name: complexa-setup
description: >
First-time setup, environment configuration, and model-weight installation for
Proteina-Complexa. Reach for this skill whenever the user says "set up complexa",
"install complexa", "configure my .env", "first-time setup", "what models do I
have installed", "what's in my .env", "download model weights", "download
Complexa / AF2 / RF3 / ProteinMPNN / LigandMPNN / ESM2 / ESMFold checkpoints",
"preflight my GPU", "verify environment", "complexa init", "complexa download",
"complexa download --status", "complexa validate env", or any time a fresh
checkout needs to be made runnable. This is the first skill to run on a new
clone — it drives `complexa init`, `complexa download`, and `complexa validate
env` end-to-end, edits the required `.env` keys, picks the right runtime (UV
vs Docker), and emits a replayable setup artifact.
compatibility: "complexa CLI installed (pip install -e .); bash 4+; nvidia-smi optional"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Complexa Setup Skill
Drive the three steps a fresh Proteina-Complexa checkout needs before any
design run: create `.env`, fetch model weights, and sanity-check the env.
Probe the host for GPU / disk / tool binaries first so the user does not
discover a missing dependency mid-pipeline. End with a JSON setup artifact the
user (or a future agent) can re-read instead of re-deriving state.
## CLI vs direct file-edit — pick the cheapest path per step
| Step | Preferred path | Why |
|---|---|---|
| `.env` creation (Step 2) | **File edit** (`cp .env_example .env` + 3 line swaps) or `complexa init` | `complexa init` is a thin wrapper around `cp + 3 regex swaps` (`_swap_runtime_in_env` in `cli_runner.py`). Either path works; pick CLI for new humans, direct edit for agents. |
| `.env` value edits (Step 3) | **File edit** (StrReplace `LOCAL_CODE_PATH=…` etc.) | No CLI for this — the values are user-specific paths. |
| Download model weights (Step 4) | **CLI** (`complexa download --…`) | Dispatches to `env/download_startup.sh` (~1000 lines of bash with NGC URLs, retries, checksum-style skip-if-present). Don't try to replicate. |
| Validate env (Step 5) | **CLI** (`complexa validate env`) or `test -f .env && test -d $DATA_PATH` | CLI prints a nicer report; the manual check is one-liner-safe. |
| Validate full design config (after picking a pipeline) | **CLI** (`complexa validate design CONFIG`) | Non-trivial Hydra defaults traversal + ckpt + env-var checks; not worth replicating. |
## What this skill enables
- A correctly-shaped `.env` for either UV or Docker runtime.
- Model checkpoints (Complexa protein/ligand/AME plus community models) downloaded to known paths.
- A `preflight.json` snapshot of the host (GPU, disk, .env, ckpts, tool binaries).
- A `run_manifest.json` capturing exactly which `complexa init` + `complexa download` invocations were used (replay-friendly).
- A pass/fail report from `complexa validate env` with clear next-step hints.
## Step 1: Pre-flight check
Always run the shared preflight before touching the environment. It does not
require `.env` to exist — it falls back to defaults — and it tells you whether
the host can run Complexa at all.
```bash
bash .claude/skills/_shared/scripts/preflight.sh
```
The script writes `./complexa_setup/preflight.json`. Read it and surface:
- `gpu.available` — if `false`, design / evaluate steps will fail; warn the user.
- `gpu.vram_gb` — Complexa needs ≥40 GB (A100/H100/L40S).
- `disk.free_gb` at `CKPT_PATH` — minimum ~50 GB for the full Complexa + community model set.
- `env.missing_required` — anything listed here must be edited in `.env` before validation passes.
- `tools.{foldseek,mmseqs,dssp,hbplus,sc}.exists` — missing tools degrade evaluation but do not block generation.
## Step 1b: Build the Python environment (only if `.venv/` is missing)
The `complexa` CLI is installed inside the project's Python environment, not on
the system path by default. On a **fresh clone**, the `.venv/` directory does
not yet exist and `complexa init` will fail with `command not found`. Build the
UV venv before anything else:
```bash
test -d .venv || ./env/build_uv_env.sh # first-time UV build
source .venv/bin/activate
which complexa # sanity check: should point inside .venv
```
Skip this step if `which complexa` already resolves — that means a previous
build is still good. The Docker runtime skips it entirely; the venv lives
inside the container image instead. If the user said "I just cloned" or you
see no `.venv/` next to `pyproject.toml`, run the build script — `complexa init`
without a venv produces a confusing `command not found` rather than an obvious
"build the venv first" error.
## Step 2: Create `.env`
Pick the runtime. UV is the default and faster to start; Docker is required on
Ubuntu 20.04 or systems with GLIBC mismatches.
Use AskUserQuestion if it is not obvious from context:
> "Which runtime do you want to configure? `uv` (recommended, faster) or `docker` (use if you do not have a UV venv built locally)?"
### Path A: file edit (preferred for agents)
`complexa init` only does three things — copy `.env_example` → `.env` and
swap the `COMPLEXA_RUNTIME=` line plus the `UV_*` ↔ `DOCKER_*` prefixes on the
tool/data/cache path block. You can do the same with `cp` + StrReplace and skip
the CLI:
```bash
cp .env_example .env
# Then StrReplace these lines in .env:
# COMPLEXA_RUNTIME=uv → COMPLEXA_RUNTIME=<runtime>
# FOLDSEEK_EXEC=${UV_FOLDSEEK_EXEC} → FOLDSEEK_EXEC=${DOCKER_FOLDSEEK_EXEC} (and same for RF3_EXEC_PATH, SC_EXEC, HBPLUS_EXEC, MMSEQS_EXEC, DSSP_EXEC, TMOL_PATH)
# DATA_PATH=${LOCAL_DATA_PATH} → DATA_PATH=${DOCKER_DATA_PATH} (and same for CACHE_DIR, CKPT_PATH)
```
Skip the prefix swap entirely if you're staying on UV (the `.env_example`
already targets UV).
### Path B: CLI
```bash
complexa init # UV runtime (default)
complexa init --runtime docker # Docker runtime
complexa init --force # Recreate .env from .env_example (drops any edits)
```
If `.env` already exists and `--force` is not passed, only the runtime-dependent
lines are swapped — user edits in Step 3 are preserved across runtime flips.
### Verify either way
```bash
test -f .env && echo "OK: .env present" || echo "MISSING"
grep -E '^COMPLEXA_RUNTIME=' .env
```
## Step 3: Edit .env
No CLI for this — Step 2 only set the runtime; you still need to write your
machine-specific paths into `.env` by hand (StrReplace or your editor). The two
absolutely-required edits are:
```bash
LOCAL_CODE_PATH=/absolute/path/to/protein-foundation-models
LOCAL_DATA_PATH=/absolute/path/to/PFM_data
```
Everything else (cache, ckpts, community-model dirs, tool binaries) is derived
from `LOCAL_CODE_PATH` by default and only needs editing if you have a
non-standard layout. For the full table — every key, what it controls, what
fails if it is missing — see [reference/env_keys.md](reference/env_keys.md).
Quick decision table for the four edits most users make:
| Key | Default | Set this if |
|-----|---------|-------------|
| `LOCAL_CODE_PATH` | placeholder | Always — required |
| `LOCAL_DATA_PATH` | `/path/to/PFM_data` | Always — required, points at target PDBs |
| `HF_TOKEN` | placeholder | You need ESMFold or gated HF models |
| `WANDB_API_KEY` | placeholder | You want training runs logged to W&B |
## Step 4: Download checkpoints
**Always use the CLI here.** `complexa download` dispatches to
`env/download_startup.sh` (~1000 lines of bash with NGC URLs, retries, and
skip-if-present logic across ~6 community-model families). Rolling your own
wget loop is a recipe for partial downloads and wrong destination paths.
Ask which models the user actually needs — downloading everything is ~100+ GB.
Pick from the three Complexa variants and the community-model set. Each
Complexa variant unlocks exactly one `complexa design` pipeline; AF2 / RF3
inside the community-model set are what `evaluate` (and reward-guided search)
need at run time.
| Flag | What it downloads | Unlocks pipeline | Destination | Approx size |
|------|-------------------|------------------|-------------|-------------|
| `--complexa` | Complexa protein-binder model + AE (`complexa.ckpt`, `complexa_ae.ckpt`) | **Protein binder** (default) — `configs/search_binder_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-ligand` | Ligand-binder model + AE (`complexa_ligand.ckpt`, `complexa_ligand_ae.ckpt`) | Ligand binder — `configs/search_ligand_binder_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-ame` | AME motif-scaffolding model + AE (`complexa_ame.ckpt`, `complexa_ame_ae.ckpt`) | AME (enzyme) — `configs/search_ame_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-all` | All three Complexa variants | All three pipelines | `./ckpts/` | ~9 GB |
| `--all` | All community models (ProteinMPNN + LigandMPNN + AF2 + ESM2 + ESMFold + RF3) | Needed by **evaluate / reward**: AF2 (protein binder), RF3 (ligand binder + AME), MPNNs (inverse folding for every pipeline). | `./community_models/` | ~50 GB |
| `--everything` | Complexa + community + optional (Boltz2 / Protenix) | Everything plus alternative refold backends | both | ~100+ GB |
| `--status` | Show install state — does not download | (none) | (none) | n/a |
**Minimum download per pipeline:**
- Protein binder (default): `complexa download --complexa --all`
- Ligand binder: `complexa download --complexa-ligand --all`
- AME / enzyme: `complexa download --complexa-ame --all`
- All three: `complexa download --everything`
For the full per-model destination breakdown and per-flag NGC sources, see
[reference/downloads.md](reference/downloads.md).
Pick the smallest invocation that covers the user's goal, then run:
```bash
complexa download --complexa # protein binder only
complexa download --complexa-all # all three Complexa variants
complexa download --all # community models only
complexa download --everything # everything
```
Without arguments, `complexa download` launches an interactive wizard — prefer
explicit flags in agent mode.
Verify what landed:
```bash
complexa download --status
```
The status output groups by family (Complexa / community / optional) and prints
"Installed" or "Missing" per ckpt. Re-run the specific flag if a ckpt is
flagged missing.
## Step 5: Validate
Final check that `.env` is loadable and the required paths resolve:
```bash
complexa validate env
```
`validate env` checks: (1) `.env` exists, (2) `DATA_PATH` is set and points at
an existing directory. It does not check ckpt files — those are checked by
`complexa validate design <config>` once you have a pipeline config picked.
Common failures and fixes:
| Symptom | Cause | Fix |
|---------|-------|-----|
| `.env file: No .env file found` | `complexa init` not run | Run `complexa init` first |
| `DATA_PATH: Not set in .env` | Placeholder not edited | Edit `LOCAL_DATA_PATH` in `.env` |
| `DATA_PATH: Directory not found` | Path edited but does not exist on disk | `mkdir -p $LOCAL_DATA_PATH` or copy target data there |
| Hydra error `InterpolationKeyError: AF2_DIR` | Reward/eval config wants AF2 but `.env` does not define it | Download AF2 weights or remove AF2 from the config |
## Step 6: Emit setup artifact
Drop a JSON manifest in `./complexa_setup/` so the user has a single file
describing the resulting state. The shared helper writes it for you:
```bash
mkdir -p ./complexa_setup
python .claude/skills/_shared/scripts/write_manifest.py \
--kind setup \
--runtime "$(grep -E '^COMPLEXA_RUNTIME=' .env | cut -d= -f2)" \
--preflight ./complexa_setup/preflight.json \
--out ./complexa_setup/run_manifest.json
```
Surface the resulting files to the user:
```bash
ls -la ./complexa_setup/
```
Expected contents:
```
complexa_setup/
├── preflight.json # GPU / disk / .env / ckpt / tool snapshot
└── run_manifest.json # init + download invocations + git SHA + runtime
```
## Hardware requirements
| Resource | Minimum | Recommended |
|----------|---------|-------------|
| GPU | 1× CUDA GPU, ≥24 GB VRAM | A100 / H100 / L40S, 40–80 GB VRAM |
| CUDA | 12.0 | 12.4+ |
| Disk (CKPT_PATH) | 50 GB | 150 GB (covers `--everything`) |
| RAM | 16 GB | 64 GB+ |
| OS | Ubuntu 22.04+ (UV) | Ubuntu 22.04+ or Docker on any host |
Ubuntu 20.04 throws GLIBC errors with the UV runtime — use `complexa init
--runtime docker` on those hosts. See `.claude/skills/_shared/reference/hardware.md`
for per-pipeline (binder vs ligand vs AME) requirements.
## Troubleshooting
| Symptom | Cause | Fix |
|---------|-------|-----|
| `complexa: command not found` | Package not installed in active env | `source .venv/bin/activate` then `pip install -e .` |
| `complexa init` says `.env_example not found` | Running outside repo root | `cd` to the project root (where `.env_example` lives) |
| `.env_example not found. Cannot initialize .env.` | Not in project root | `cd` into `protein-foundation-models/` and retry |
| `complexa download` fails on NGC URL | Behind firewall / no internet | Configure a proxy for the download script, or download the model `.ckpt`s manually from the NGC pages linked in the main `README.md` and drop them into `./ckpts/` |
| `complexa download --status` shows ckpts present but `validate` fails | `.env` `CKPT_PATH` points elsewhere | Either move ckpts or edit `LOCAL_CHECKPOINT_PATH` in `.env` |
| GLIBC error on import | Ubuntu 20.04 with UV runtime | Re-run `complexa init --runtime docker` and use `./env/docker-ops.sh run` |
---
For the full `.env` reference (every key, defaults, failure modes), see
[reference/env_keys.md](reference/env_keys.md).
For the full download flag matrix, NGC URLs, and destination layout, see
[reference/downloads.md](reference/downloads.md).
Referenced files: 2
complexa-sweep12.4 KB
---
name: complexa-sweep
description: Use this skill whenever the user wants to run a parameter sweep over a Proteina-Complexa design pipeline — cartesian-product hyperparameter scans, Pareto search over generation/reward/evaluation knobs, or any "compare configurations" workflow. Trigger phrases include "sweep beam width", "sweep nsteps", "hyperparameter sweep", "parameter scan", "scan beam_width and temperature", "compare configurations", "find the best generation params", "what's the optimal nsteps", "Pareto search for binder quality vs wall-clock", "complexa sweep", "tune Complexa", "ablate the reward weights", "configs/sweeps", "--sweeper", "run beam_width.yaml". This is the only skill that owns sweeper YAML authoring, cartesian-product expansion, and per-config result ranking.
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# complexa-sweep
Run cartesian-product parameter sweeps over Proteina-Complexa design pipelines. Pick or author a sweeper YAML in `configs/sweeps/`, expand it to N inference configs with `script_utils/generate_inference_configs.py`, loop `complexa design` over those configs, then aggregate per-config success metrics into a ranked summary CSV plus a manifest.
> **Important:** The `complexa design` CLI does **NOT** accept `--sweeper` directly. Sweeps are driven by `script_utils/generate_inference_configs.py`, which writes one `configs/inference_configs/inf_{idx}_{run_name}.yaml` per sweep combination. You then loop `complexa design` over those generated configs.
## What this skill enables
- Pick an existing sweeper YAML from `configs/sweeps/` (`beam_width`, `bb_ca_temperature`, `search_replicas`, `example`).
- Author a new sweeper YAML with arbitrary dot-notation axes (cartesian product).
- Generate N inference + evaluation config pairs with `script_utils/generate_inference_configs.py`.
- Loop `complexa design` over the generated configs (one at a time on a single GPU, or in parallel on a multi-GPU host by sharding the config list across `CUDA_VISIBLE_DEVICES`).
- Walk per-config output directories and parse the analyze-step CSV from each.
- Emit `sweep_summary.csv` (one row per config: axis values + success rate + mean iPAE + diversity) and `sweep_manifest.json`.
- Identify the best config by success rate and the Pareto frontier (wall-clock vs success).
## Step 1: Pre-flight
```bash
bash .claude/skills/_shared/scripts/preflight.sh
```
Read `./complexa_setup/preflight.json`. A sweep multiplies GPU time by the number of configs. **Before launching, confirm the cost with the user**:
> "This sweep produces N configs × ~M minutes per config ≈ TOTAL GPU-hours. OK to proceed? (y / reduce / cancel)"
If `gpu.available=false`, stop — sweeps are not feasible on CPU.
## Step 2: Pick the pipeline + target
Use the same dialogue as `complexa-design` — do **not** duplicate it here. See [`.claude/skills/complexa-design/SKILL.md`](../complexa-design/SKILL.md) Step 2 ("Pick the pipeline") and Step 3 ("Gather parameters"). Capture:
- `pipeline_config_name` — e.g. `search_binder_local_pipeline` (default), `search_ligand_binder_local_pipeline`, `search_ame_local_pipeline`.
- `task_name` — e.g. `02_PDL1`, `22_DerF21`, `39_7V11_LIGAND`. Passed as `--override generation.task_name=<task>`.
- `run_name` — short tag for output dir naming.
## Step 3: Pick or author the sweeper YAML
Sweeper YAMLs live in `configs/sweeps/`. Each key is a dot-notation Hydra path; each value is a list. The cartesian product becomes N configs.
### Canned sweepers
| File | Axis | Values | Configs |
|---|---|---|---|
| `configs/sweeps/beam_width.yaml` | `generation.search.beam_search.beam_width` | 1, 2, 4, 8 | 4 |
| `configs/sweeps/bb_ca_temperature.yaml` | `generation.model.bb_ca.simulation_step_params.sc_scale_noise` | 0.1, 0.4 | 2 |
| `configs/sweeps/search_replicas.yaml` | `generation.search.best_of_n.replicas` | 1, 4, 16, 64 | 4 |
| `configs/sweeps/example.yaml` | beam_width × nsteps | (2,4) × (200,400) | 4 |
If one matches the user's intent, use it as-is. Otherwise author a new file.
### Authoring a new sweeper
Minimal multi-axis example (saved to `configs/sweeps/my_sweep.yaml`):
```yaml
# 3 beam widths × 2 nsteps = 6 configs
generation.search.beam_search.beam_width:
- 2
- 4
- 8
generation.args.nsteps:
- 200
- 400
```
Rules (from `script_utils/generate_inference_configs.py:load_sweeper_file`):
- Top-level mapping only. Keys are dot-notation Hydra paths.
- Values must be **lists**. A scalar is auto-wrapped into a single-element list (which pins a value without adding a dimension).
- Cartesian product: total configs = product of list lengths. Two 4-value axes = 16 configs; budget accordingly.
- If a key appears in both the sweeper file and an `--override`, the override wins and that axis collapses.
See [reference/sweep_axes.md](reference/sweep_axes.md) for the full catalogue of swept keys (typical ranges, cost multipliers, what improves/regresses).
### Dry-run preview before generating
Always confirm the config count first:
```bash
python script_utils/generate_inference_configs.py \
--config_name search_binder_local_pipeline \
--sweeper configs/sweeps/my_sweep.yaml \
--override generation.task_name=22_DerF21 \
--run_name my_sweep \
--dryrun
```
The output lists every axis + value list and prints `DRY RUN — would generate N config pair(s)`.
## Step 4: Generate configs + loop `complexa design`
Once the dry-run looks right, drop `--dryrun` to materialize `inf_{idx}_{run_name}.yaml` + `eval_{idx}_{run_name}.yaml` pairs under `configs/inference_configs/` and `configs/eval_configs/`. Then loop `complexa design` over the inference configs.
```bash
# 1. Generate one inf_*.yaml + one eval_*.yaml per combination.
python script_utils/generate_inference_configs.py \
--config_name search_binder_local_pipeline \
--sweeper configs/sweeps/my_sweep.yaml \
--override generation.task_name=22_DerF21 \
--run_name my_sweep
# 2. Loop complexa design over each generated inference config.
for cfg in configs/inference_configs/inf_*_my_sweep.yaml; do
complexa design "$cfg" || echo "FAILED: $cfg"
done
```
This serialises the sweep on a single GPU — be honest with the user that wall-clock = N × per-run.
### Multi-GPU host (optional speed-up)
On a host with K GPUs, shard the config list K ways and launch one loop per GPU in parallel. Each `complexa design` call uses exactly one GPU at default `gen_njobs=1` / `eval_njobs=1`, so pinning via `CUDA_VISIBLE_DEVICES` keeps them from colliding:
```bash
CONFIGS=(configs/inference_configs/inf_*_my_sweep.yaml)
N=${#CONFIGS[@]}; K=4 # 4 GPUs
for gpu in $(seq 0 $((K-1))); do
(
for i in $(seq $gpu $K $((N-1))); do
CUDA_VISIBLE_DEVICES=$gpu complexa design "${CONFIGS[$i]}" \
|| echo "FAILED on GPU $gpu: ${CONFIGS[$i]}"
done
) &
done
wait
```
Drop `gen_njobs` / `eval_njobs` overrides into the loop only if you intentionally want each `complexa design` invocation to consume multiple GPUs (and you have a way to keep them out of each other's way).
## Step 5: Collect results
Each swept config writes to its own `./inference/inf_{idx}_{run_name}/` directory; the analyze step writes `./evaluation_results/eval_{idx}_{run_name}/results_*.csv`. After the sweep finishes:
```bash
ls -d ./inference/inf_*_my_sweep/
ls ./evaluation_results/eval_*_my_sweep/results_*.csv
```
For each config, parse the analyze CSV (one row per generated binder). Standard columns used for ranking: `i_pae`, `i_plddt`, `sc_rmsd`, `binder_seq`, `passes_filter` (bool). If `passes_filter` is missing, derive a success flag with the user's chosen thresholds (defaults: `i_pae < 10`, `i_plddt > 0.7`, `sc_rmsd < 2.0` — confirm with user).
## Step 6: Rank configs
Emit `sweep_summary.csv` to the run directory. One row per config:
| Column | How to compute |
|---|---|
| `config_id` | The `{idx}` from `inf_{idx}_{run_name}` |
| `<axis_1>`, `<axis_2>`, ... | The swept value at this combination (read from the per-config `inf_*.yaml`) |
| `n_samples` | Row count in the analyze CSV |
| `success_rate` | `passes_filter.mean()` |
| `mean_i_pae` | `i_pae.mean()` (lower = better) |
| `mean_i_plddt` | `i_plddt.mean()` (higher = better) |
| `diversity_score` | Unique sequence count / `n_samples` (or use TM-score clustering if available) |
| `wall_clock_min` | From the per-config log timestamps |
Then report:
- **Best config** = argmax of `success_rate`. Tie-break on `mean_i_pae` (lower).
- **Pareto frontier** over (`wall_clock_min`, `success_rate`): a config is on the frontier iff no other config is both faster AND has higher success rate.
Print the best config + the frontier to the terminal. Save the full table to `sweep_summary.csv`.
## Step 7: Emit manifest
Capture the resolved invocation + outputs for replay. The shared helper takes a single `--output-dir` (it walks for CSVs and pulls the Hydra config from the run's `.hydra/`), so call it against the best-config's evaluation directory and pass the parent `complexa design` command for that run:
```bash
BEST_ID=4 # from Step 6 ranking
python3 .claude/skills/_shared/scripts/write_manifest.py \
--output-dir ./evaluation_results/eval_${BEST_ID}_my_sweep \
--command "python script_utils/generate_inference_configs.py --config_name search_binder_local_pipeline --sweeper configs/sweeps/my_sweep.yaml --override generation.task_name=22_DerF21 --run_name my_sweep && for cfg in configs/inference_configs/inf_*_my_sweep.yaml; do complexa design \"\$cfg\"; done" \
--skill complexa-sweep \
--out ./sweep_runs/my_sweep/sweep_manifest.json
```
Alongside the manifest, save the ranked `sweep_summary.csv` from Step 6 to `./sweep_runs/my_sweep/sweep_summary.csv` and surface both paths to the user.
## Recommended sweep recipes
| Symptom | Sweep this | Why |
|---|---|---|
| Quality not good enough | `beam_width × nsteps` (use `example.yaml` as a starting point) | Both raise compute → quality; find the cheapest combination that lands. |
| Too slow / want speed-up | `nsteps` downward (e.g. `[100, 200, 400]`) with fixed `beam_width=4` | Find the smallest nsteps that retains success rate. |
| Mode collapse / low diversity | `generation.model.bb_ca.simulation_step_params.sc_scale_noise` (use `bb_ca_temperature.yaml`) | Higher noise → more diverse backbones. |
| Reward over-fitting (high reward, bad metrics) | `generation.reward_model.reward_models.af2folding.reward_weights.{i_pae, plddt}` ratio | Re-balance composite reward. |
| Want statistical robustness on one config | `search_replicas.yaml` (`best_of_n.replicas`) | Same config, more samples → tighter success-rate estimate. |
| Algorithm shoot-out | `generation.search.algorithm` over `[single-pass, best-of-n, beam-search, fk-steering]` | Compare search regimes; pin everything else. |
## Hardware
Total GPU-time for a sweep = `N_configs × per_run_GPU_time`. A single `search_binder_local_pipeline` run on one A100 is roughly 30–90 min (binder length + nsteps dependent). A 4-axis × 4-value sweep = 256 configs × ~60 min = ~256 GPU-hours — at that scale, plan on a multi-GPU host with the Step-4 sharding pattern.
Refer to [`.claude/skills/_shared/reference/hardware.md`](../_shared/reference/hardware.md) for the per-run baseline + VRAM minima.
## Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| `Sweeper file not found` from `generate_inference_configs.py` | Path resolved from wrong CWD | Use a path relative to the repo root; or pass an absolute path. |
| `Sweeper YAML must be a mapping` | Top-level YAML is a list or scalar | Rewrite as `key: [v1, v2]` mapping. |
| `No configs were generated` | One of the value lists is empty `[]` | Sweeper file has a `key: []` line — add at least one value. |
| `complexa design` rejects `--sweeper` | Confusion between entrypoints | The CLI does not accept `--sweeper`. Use `generate_inference_configs.py` first, then loop `complexa design` over the generated configs. |
| Override silently collapses a sweep axis | `--override key=v` shadowed a sweep key | Drop either the override OR the matching key from the sweeper file. |
| One config in the loop fails, sweep keeps going | The `\|\| echo "FAILED…"` in Step 4 swallows the error | Re-run the failed `inf_*.yaml` standalone; check the per-config log under `./logs/`. Skip failed `config_id` when ranking. |
---
For per-axis reference (typical ranges, cost, what gets better/worse), see [reference/sweep_axes.md](reference/sweep_axes.md).
For the user-facing sweep system overview (config generation, output layout), see [`docs/SWEEP.md`](reference/SWEEP.md).
Referenced files: 2
complexa-target14.6 KB
---
name: complexa-target
description: Use this skill whenever the user wants to add, register, edit, list, show, or validate a Proteina-Complexa design target for any pipeline — protein binder (default), ligand binder, or AME / enzyme scaffolding. Triggers include "add a target", "define a new target for binder design", "register a hotspot", "set up a PDL1 binder target", "ligand binder pocket", "SMILES target", "AME task", "enzyme motif", "M0024_1nzy", "complexa target add", "complexa target show", "configure target X", "what targets are available", "where do hotspots live", "what does target_input mean", "chain-spec syntax", "binder length range", or any question about `configs/targets/{,ligand_}targets_dict.yaml` and `configs/design_tasks/ame_dict_v2.yaml`. Also covers `complexa validate target`. This is the only skill that touches the three targets dict files.
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# complexa-target
Add or edit a design target in Proteina-Complexa. Targets live in **three YAML files**, one per `complexa design` pipeline:
- `configs/targets/targets_dict.yaml` — protein binder (**default pipeline**)
- `configs/targets/ligand_targets_dict.yaml` — ligand binder
- `configs/design_tasks/ame_dict_v2.yaml` — AME / enzyme scaffolding
The `complexa target` CLI only manages the first two. AME tasks use an extended schema (the same core fields as ligand targets plus `contig_atoms` for hand-curated per-residue motif atom selections) and are file-edit-only.
## Preferred path: edit the YAML directly
`complexa target add` is a thin wrapper around "load YAML → build dict → append a block" (see `src/proteinfoundation/cli/target_manager.py:add_target_cli` and `append_target_to_dict`). For agentic use, **just edit the targets dict directly** — the schema is short, the existing entries are great copy templates, and you skip 14 CLI flags / SMILES shell-escaping. Use the CLI only if you explicitly want its YAML auto-quoter or its overwrite-prompt safety.
The skill therefore presents both paths:
| Step | Direct file edit (preferred) | CLI |
|---|---|---|
| Look at existing targets | Read `configs/targets/{,ligand_}targets_dict.yaml` (or `configs/design_tasks/ame_dict_v2.yaml` for AME) | `complexa target list` (protein/ligand only) |
| Look at one target | Read the dict and grep for the key | `complexa target show NAME` |
| Add a new target | Append a YAML block (Step 3a) | `complexa target add ...` (Step 3b) |
| Verify the PDB resolves on disk | — *(no shortcut, this checks Hydra defaults)* | `complexa validate target CONFIG --target NAME` |
`complexa target` only sees the two protein-style dicts (`targets_dict.yaml` and `ligand_targets_dict.yaml`). For AME tasks (which use a different schema), file-edit is the only path — see Step 1 below.
## What this skill enables
- Register a **protein target** (chain + residue range + hotspots) in `configs/targets/targets_dict.yaml`.
- Register a **ligand target** (PDB pocket + 3-letter code + SMILES) in `configs/targets/ligand_targets_dict.yaml`.
- Resolve the right **AME task name** (e.g. `M0024_1nzy`, `M0096_1chm`) for the AME pipeline — these are not added via `complexa target add` (see Step 1).
- Verify a target resolves to a real PDB on disk with `complexa validate target`.
- Emit a replayable artifact (`target_definition.yaml`) for downstream design runs.
## Step 1: Decide the target type
Targets live in **three different files**, one per `complexa design` pipeline. Pick the file by working out which pipeline the user will run, then add the entry to that file. Each file is consumed by exactly one pipeline.
| User intent | Pipeline | Targets dict file | This skill? |
|---|---|---|---|
| Bind a protein surface (PD-L1, IFNAR2, TNF-α, …) — **default** | `configs/search_binder_local_pipeline.yaml` | `configs/targets/targets_dict.yaml` | Yes — Step 2 (protein) |
| Bind a small-molecule pocket (FAD, SAM, OQO, …) | `configs/search_ligand_binder_local_pipeline.yaml` | `configs/targets/ligand_targets_dict.yaml` | Yes — Step 2 (ligand) |
| AME / enzyme scaffolding (motif + ligand, `M####_<pdb>` names) | `configs/search_ame_local_pipeline.yaml` | `configs/design_tasks/ame_dict_v2.yaml` | Partial — see "AME tasks" below |
### AME tasks (enzyme pipeline)
Similar story — AME tasks live in `configs/design_tasks/ame_dict_v2.yaml` under `motif_target_dict_cfg:` with their own schema (`source`, `target_filename`, `ligand`, `contig_atoms`, `binder_length`, `use_bonds_from_file`). The `contig_atoms` string encodes per-residue motif atom selections like `"A64: [O, C]; A86: [CB, CA, N, C]; ..."` — these are hand-curated, so adding a new AME task is a file edit, not a CLI invocation. Browse the file, copy a similar entry as a template, and pass `++generation.task_name=<NAME>` to `complexa design configs/search_ame_local_pipeline.yaml`.
## Step 2: Gather required info
Use AskUserQuestion to collect — do not guess these. Required fields differ for protein vs ligand.
### Protein target
| Field | Question | Example | Required |
|---|---|---|---|
| name | "Target name (used as the dict key and `task_name`)?" | `02_PDL1`, `MyTarget_v1` | yes |
| source | "Source directory under `$DATA_PATH/target_data/`?" | `bindcraft_targets`, `custom_targets` | yes (or `target_path`) |
| target_filename | "PDB filename (no `.pdb` extension)?" | `PD-L1`, `IFNAR2` | yes (or `target_path`) |
| target_input | "Chain + residue range — see reference for grammar." | `A1-115`, `A1-50,B1-50` | yes |
| hotspot_residues | "Hotspot residues (interface contact residues)?" | `["A33", "A95", "A102"]` | optional, recommended |
| binder_length | "Binder length range `[min, max]` or single `[length]`?" | `[80, 150]`, `[100]` | optional (default `[60, 120]`) |
| pdb_id | "Reference PDB ID (optional, metadata only)?" | `"2lag"` | optional |
### Ligand target — protein fields above (minus `target_input`), plus:
| Field | Question | Example | Required |
|---|---|---|---|
| ligand | "3-letter PDB ligand residue code?" | `FAD`, `OQO`, `SAM` | yes (presence marks target as ligand) |
| smiles | "SMILES string for the ligand?" | `"O=C2C3=Nc1cc(c(...)..."` | recommended |
| ligand_only | "Generate pocket around ligand only (no protein-protein interface)?" | `true` / `false` | optional (default `true`) |
| use_bonds_from_file | "Use bond info from the input PDB/CIF?" | `true` / `false` | optional (default `true`) |
| target_input | not required for ligand targets | — | no |
Check for name collisions by reading the dict (`rg '^ NEW_NAME:' configs/targets/targets_dict.yaml`) or running `complexa target list -v --ligand` / `--protein`.
## Step 3a: Append the YAML block directly (preferred)
Open `configs/targets/targets_dict.yaml` (or `ligand_targets_dict.yaml` for ligand targets), find a similar existing entry as a style template, and append the new block under `target_dict_cfg:`. Two-space indent, single blank line between entries.
### Protein template
```yaml
02_PDL1:
source: bindcraft_targets
target_filename: PD-L1
target_input: "A1-115"
hotspot_residues: ["A37", "A39", "A49", "A98"]
binder_length: [64, 155]
pdb_id: "4z18"
```
Rules to match the existing file style (mirrors what `complexa target add` would emit):
- Quote chain/residue ranges (`"A1-115"`, `"A33"`) and any SMILES — flow-style strings get tripped up by `:` and brackets otherwise.
- Use flow-style lists (`["A33", "A95"]`, `[64, 155]`) — that's what the on-disk dump produces.
- `target_input`, `source`, `target_filename` are required for protein. `target_path` (absolute path) can replace `source + target_filename` if the PDB lives outside `$DATA_PATH/target_data/`.
### Ligand template
```yaml
41_7BKC_LIGAND:
source: ligand_targets
target_filename: 7BKC_ligand_centered
hotspot_residues: [null]
binder_length: [100]
pdb_id: 7BKC
ligand: 'FAD'
ligand_only: True
SMILES: "O=C2C3=Nc1cc(c(cc1N(C3=NC(=O)N2)CC(O)C(O)C(O)COP(=O)(O)OP(=O)(O)OCC6OC(n5cnc4c(ncnc45)N)C(O)C6O)C)C"
use_bonds_from_file: True
```
The presence of the `ligand:` key flips the target into ligand mode — there is no separate `is_ligand` flag. `target_input` is not required for ligands.
After saving, skip to Step 4 to verify.
## Step 3b: CLI alternative (`complexa target add`)
Use the CLI when you want its automatic chain/residue quoting, the overwrite-confirm prompt, or to wire target creation into a non-Python script. Prefer non-interactive mode for agentic use (`-i / --editor` opens an editor and blocks).
### Confirmed flags (from `src/proteinfoundation/cli/target_cli.py`)
| Flag | Type | Applies to | Notes |
|---|---|---|---|
| `name` (positional) | str | both | dict key |
| `--dict PATH` | path | both | override target dict path |
| `-i, --interactive` | flag | both | open `$EDITOR` |
| `-e, --editor NAME` | str | both | `code`, `nano`, `vim`, `cursor`, ... |
| `--source NAME` | str | both | directory under `$DATA_PATH/target_data/` |
| `--target-filename NAME` | str | both | PDB stem (no `.pdb`) |
| `--target-path PATH` | str | both | full path; overrides `source`+`filename` |
| `--target-input SPEC` | str | protein | chain/residue range, e.g. `A1-115` |
| `--hotspot-residues R [R ...]` | list | both | e.g. `A33 A95` |
| `--binder-length N [N ...]` | int list | both | `[min, max]` or `[length]` |
| `--pdb-id ID` | str | both | metadata only |
| `--ligand CODE` | str | ligand | presence marks ligand target |
| `--ligand-only` | flag | ligand | pocket-only mode |
| `--smiles STR` | str | ligand | SMILES for the ligand |
| `--use-bonds-from-file` | flag | ligand | use PDB bonds |
| `-f, --force` | flag | both | overwrite existing without prompt |
### Protein example
```bash
complexa target add 02_PDL1 \
--source bindcraft_targets \
--target-filename PD-L1 \
--target-input A1-115 \
--hotspot-residues A37 A39 A49 A98 \
--binder-length 64 155 \
--pdb-id 4z18
```
### Ligand example
```bash
complexa target add 41_7BKC_LIGAND \
--source ligand_targets \
--target-filename 7BKC_ligand_centered \
--pdb-id 7BKC \
--ligand FAD \
--ligand-only \
--use-bonds-from-file \
--smiles "O=C2C3=Nc1cc(c(cc1N(C3=NC(=O)N2)CC(O)C(O)C(O)COP(=O)(O)OP(=O)(O)OCC6OC(n5cnc4c(ncnc45)N)C(O)C6O)C)C" \
--binder-length 100
```
Why these defaults: `--source` defaults to `custom_targets` (protein) or `ligand_targets` (ligand) and `--target-filename` defaults to `name`, so be explicit when the PDB stem differs from the target name. The dict gets `ligand_only: true` and `use_bonds_from_file: true` by default for ligand targets — pass the flags to be explicit, or omit to accept defaults.
## Step 4: Verify
Run two checks, in order. Stop and fix on the first failure.
```bash
# 1. Confirm the entry landed and looks right.
# Direct read works just as well: `rg -A 7 '^ 02_PDL1:' configs/targets/targets_dict.yaml`
complexa target show 02_PDL1
# 2. Resolve the PDB path and validate it exists on disk.
# No shortcut for this — it traverses Hydra defaults to find the target dict.
complexa validate target configs/search_protein_local_pipeline.yaml --target 02_PDL1
```
`complexa validate target` (from `src/proteinfoundation/cli/validate.py::validate_target`) checks:
- `DATA_PATH` is set and `target_data/` exists.
- The target name resolves in `target_dict_cfg`.
- Either `target_path` exists, or `$DATA_PATH/target_data/<source>/<target_filename>.pdb` exists.
- `target_input`, `hotspot_residues`, `binder_length` are reported back so a human can sanity-check them.
For ligand targets, point `--target` at a ligand pipeline config (e.g. a ligand binder search config).
## Step 5: Emit artifact
Save a replayable record under `./target_<name>/`:
```bash
mkdir -p target_02_PDL1
complexa target show 02_PDL1 > target_02_PDL1/target_show.txt
```
Then write the appended YAML snippet (the lines `complexa target add` wrote under `target_dict_cfg:`) to `target_02_PDL1/target_definition.yaml` using the Write tool. The artifact lets the user diff target definitions across runs and re-create the entry on another checkout via `complexa target add ... --force`.
## Hardware requirements
None for target definition — this is a YAML edit, not a training/inference step. Disk impact is a few KB appended to the targets dict (plus an automatic `.yaml.bak` backup written by `save_targets_dict`).
For the downstream design / evaluate runs that consume the target, defer to `complexa-design` and `_shared/reference/hardware.md`.
## Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| `Target PDB file: File not found` from `validate target` | `<source>/<target_filename>.pdb` does not exist under `$DATA_PATH/target_data/` | Confirm the PDB stem and source dir. Use `--target-path /full/path.pdb` if the file lives outside `target_data/`. |
| "Target 'X' already exists! Overwrite? (y/N)" | Name collision in dict | Either pick a new name, or pass `-f / --force` to overwrite. |
| Hotspot residue not in PDB | Wrong chain or residue number | Open the PDB, re-check chain letters (case-sensitive) and residue indices. Hotspots use the format `<CHAIN><RESNUM>` — see reference. |
| Chain not found | `target_input` references a chain that does not exist in the PDB | Inspect the PDB with `grep "^ATOM" target.pdb \| awk '{print $5}' \| sort -u`. |
| Ligand code missing from PDB | The 3-letter `ligand` code does not appear as a `HETATM` residue name in the file | Open the PDB and check `HETATM` lines; you may need a `_ligand_centered` variant of the PDB. |
| SMILES parse failure downstream | Bad SMILES string | Validate with `rdkit.Chem.MolFromSmiles(smiles)` before adding. Quote SMILES in shell to escape brackets and parens. |
| Target name with leading digit breaks Hydra interpolation | YAML treats `02_PDL1` as a string fine; some override syntaxes need quoting | Use `++generation.task_name=02_PDL1`; quote if the shell strips characters. |
| `target_input` appears ignored for a ligand target | By design — ligand targets do not use `target_input`; pocket is defined by the ligand | Leave it unset for ligand targets. |
## Reference
- `reference/target_schema.md` — every field, chain-spec grammar, AME task-name grammar, three worked examples.
- `configs/targets/targets_dict.yaml` — live protein entries (copy a known-good one as a template).
- `configs/targets/ligand_targets_dict.yaml` — live ligand entries.
- `configs/design_tasks/ame_dict_v2.yaml` — AME task definitions (file-edit only, not exposed via `complexa target` CLI).
- `src/proteinfoundation/cli/target_cli.py` — argparse source of truth.
- `src/proteinfoundation/cli/target_manager.py` — `add_target_cli`, `list_targets`, `show_target`, schema in `TARGET_FIELDS`.
- `src/proteinfoundation/cli/validate.py` — `validate_target` implementation.
Referenced files: 1
cuequivariance15.3 KB
---
name: cuequivariance
description: Define custom groups (Irrep subclasses), build segmented tensor products with CG coefficients, create equivariant polynomials and IrDictPolynomials, and use built-in descriptors (linear, tensor products, spherical harmonics). Use when working with cuequivariance group theory, irreps, or segmented polynomials.
---
# cuequivariance: Groups, Irreps, and Segmented Polynomials
## Overview
`cuequivariance` (imported as `cue`) provides two core abstractions:
1. **Group theory**: `Irrep` subclasses define irreducible representations of Lie groups (SO3, O3, SU2, or custom). `Irreps` manages collections with multiplicities.
2. **Segmented polynomials**: `SegmentedTensorProduct` describes tensor contractions over segments of varying shape, linked by `Path` objects carrying Clebsch-Gordan coefficients. `SegmentedPolynomial` wraps multiple STPs into a polynomial with named inputs/outputs. Two higher-level wrappers attach group representations:
- `EquivariantPolynomial` — dense operands with `IrrepsAndLayout` metadata
- `IrDictPolynomial` — operands already split by irrep, with per-group `Irreps` metadata for the `dict[Irrep, Array]` workflow
## Defining a custom group
Subclass `cue.Irrep` (a frozen dataclass) and implement:
```python
from __future__ import annotations
import dataclasses
import re
from typing import Iterator
import numpy as np
import cuequivariance as cue
@dataclasses.dataclass(frozen=True)
class Z2(cue.Irrep):
odd: bool # dataclass field -- required for correct __eq__ and __hash__
# No __init__ needed -- @dataclass(frozen=True) generates it: Z2(odd=True)
@classmethod
def regexp_pattern(cls) -> re.Pattern:
# Pattern whose first group is passed to from_string
return re.compile(r"(odd|even)")
@classmethod
def from_string(cls, string: str) -> Z2:
return cls(odd=string == "odd")
def __repr__(rep: Z2) -> str:
return "odd" if rep.odd else "even"
def __mul__(rep1: Z2, rep2: Z2) -> Iterator[Z2]:
# Selection rule: which irreps appear in the tensor product rep1 x rep2
return [Z2(odd=rep1.odd ^ rep2.odd)]
@classmethod
def clebsch_gordan(cls, rep1: Z2, rep2: Z2, rep3: Z2) -> np.ndarray:
# Shape: (num_paths, rep1.dim, rep2.dim, rep3.dim)
if rep3 in rep1 * rep2:
return np.array([[[[1]]]])
else:
return np.zeros((0, 1, 1, 1))
@property
def dim(rep: Z2) -> int:
return 1
def __lt__(rep1: Z2, rep2: Z2) -> bool:
# Ordering for sorting; dimension is compared first by the base class
return rep1.odd < rep2.odd
@classmethod
def iterator(cls) -> Iterator[Z2]:
# Must yield trivial irrep first
for odd in [False, True]:
yield Z2(odd=odd)
def discrete_generators(rep: Z2) -> np.ndarray:
# Shape: (num_generators, dim, dim)
if rep.odd:
return -np.ones((1, 1, 1))
else:
return np.ones((1, 1, 1))
def continuous_generators(rep: Z2) -> np.ndarray:
# Shape: (lie_dim, dim, dim) -- Z2 is discrete, so lie_dim=0
return np.zeros((0, rep.dim, rep.dim))
def algebra(self) -> np.ndarray:
# Shape: (lie_dim, lie_dim, lie_dim) -- structure constants [X_i, X_j] = A_ijk X_k
return np.zeros((0, 0, 0))
# Usage:
irreps = cue.Irreps(Z2, "3x odd + 2x even") # dim=5
```
### Required methods summary
| Method | Returns | Purpose |
|--------|---------|---------|
| `regexp_pattern()` | `re.Pattern` | Parse string like `"1"`, `"0e"`, `"odd"` |
| `from_string(s)` | `Irrep` | Construct from matched string |
| `__repr__` | `str` | Canonical string form |
| `__mul__(a, b)` | `Iterator[Irrep]` | Selection rule for tensor product |
| `clebsch_gordan(a, b, c)` | `ndarray (n, d1, d2, d3)` | CG coefficients |
| `dim` (property) | `int` | Dimension of representation |
| `__lt__(a, b)` | `bool` | Ordering (dimension first, then custom) |
| `iterator()` | `Iterator[Irrep]` | All irreps, trivial first |
| `continuous_generators()` | `ndarray (lie_dim, dim, dim)` | Lie algebra generators |
| `discrete_generators()` | `ndarray (n, dim, dim)` | Finite symmetry generators |
| `algebra()` | `ndarray (lie_dim, lie_dim, lie_dim)` | Structure constants |
### Built-in groups
- **`cue.SO3(l)`**: 3D rotations. `l` is a non-negative integer. `dim = 2l+1`. String: `"0"`, `"1"`, `"2"`.
- **`cue.O3(l, p)`**: 3D rotations + parity. `p=+1` (even) or `p=-1` (odd). String: `"0e"`, `"1o"`, `"2e"`.
- **`cue.SU2(j)`**: Spin group. `j` is a non-negative half-integer. String: `"0"`, `"1/2"`, `"1"`.
## Irreps and layout
```python
irreps = cue.Irreps("SO3", "16x0 + 4x1 + 2x2") # 16 scalars, 4 vectors, 2 rank-2
irreps.dim # 16*1 + 4*3 + 2*5 = 38
for mul, ir in irreps:
print(mul, ir, ir.dim) # 16 0 1, then 4 1 3, then 2 2 5
```
`IrrepsLayout` controls memory order within each `(mul, ir)` block:
- `cue.ir_mul`: data ordered as `(ir.dim, mul)` — **used by all descriptors and ir_dict internally**
- `cue.mul_ir`: data ordered as `(mul, ir.dim)` — **used by nnx `dict[Irrep, Array]` and PyTorch**
`IrrepsAndLayout` combines irreps with a layout into a `Rep`:
```python
rep = cue.IrrepsAndLayout(cue.Irreps("SO3", "4x0 + 2x1"), cue.ir_mul)
rep.dim # 10
```
## Building a SegmentedTensorProduct from scratch
The subscripts string uses Einstein notation. Operands are comma-separated, coefficient modes follow `+`.
```python
# Matrix-vector multiply: y_i = sum_j M_ij * x_j
d = cue.SegmentedTensorProduct.from_subscripts("ij,j,i")
d.add_segment(0, (3, 4)) # operand 0: matrix segment of shape (3, 4)
d.add_segment(1, (4,)) # operand 1: vector of size 4
d.add_segment(2, (3,)) # operand 2: output vector of size 3
d.add_path(0, 0, 0, c=1.0) # link segments 0,0,0 with coefficient=1.0
poly = cue.SegmentedPolynomial.eval_last_operand(d) # last operand becomes output
[y] = poly(M_flat, x) # numpy evaluation
```
### Multi-segment STP (how descriptors work internally)
Descriptors build STPs with multiple segments per operand. Each segment corresponds to an irrep block:
```python
# Linear equivariant map: output[iv] = sum_u weight[uv] * input[iu]
d = cue.SegmentedTensorProduct.from_subscripts("uv,iu,iv")
# Segment for l=1: ir_dim=3, mul_in=2, mul_out=5
s_in_0 = d.add_segment(1, (3, 2)) # input block
s_out_0 = d.add_segment(2, (3, 5)) # output block
d.add_path((2, 5), s_in_0, s_out_0, c=1.0)
# Segment for l=0: ir_dim=1, mul_in=4, mul_out=3
s_in_1 = d.add_segment(1, (1, 4))
s_out_1 = d.add_segment(2, (1, 3))
d.add_path((4, 3), s_in_1, s_out_1, c=1.0)
```
### Weights operand
For weighted tensor products (subscript starting with `uvw` or `uv`), the first operand is always weights. The weight segment shape is `(mul_1, mul_2, ...)` matching the multiplicity modes. The weights operand gets `new_scalars()` irreps since weights are invariant.
### CG coefficients as path coefficients
```python
d = cue.SegmentedTensorProduct.from_subscripts("uvw,iu,jv,kw+ijk")
# For each pair of input irreps and each output irrep in the selection rule:
for cg in cue.clebsch_gordan(ir1, ir2, ir3):
# cg has shape (ir1.dim, ir2.dim, ir3.dim)
d.add_path((mul1, mul2, mul3), seg_in1, seg_in2, seg_out, c=cg)
```
## Descriptors
All descriptors come in two variants:
- **Original** — returns `EquivariantPolynomial` with dense operands
- **`_ir_dict`** — returns `IrDictPolynomial` with operands already split by irrep
### EquivariantPolynomial descriptors
```python
# Fully connected tensor product (all input-output irrep combinations)
e = cue.descriptors.fully_connected_tensor_product(
16 * cue.Irreps("SO3", "0 + 1 + 2"),
16 * cue.Irreps("SO3", "0 + 1 + 2"),
16 * cue.Irreps("SO3", "0 + 1 + 2"),
)
# Channelwise tensor product (same-channel only, sparse)
e = cue.descriptors.channelwise_tensor_product(
64 * cue.Irreps("SO3", "0 + 1"), cue.Irreps("SO3", "0 + 1"),
cue.Irreps("SO3", "0 + 1"), simplify_irreps3=True,
)
# Full (weightless) tensor product
e = cue.descriptors.full_tensor_product(
cue.Irreps("SO3", "2x0 + 1x1"), cue.Irreps("SO3", "0 + 1"),
)
# Elementwise tensor product (paired channels)
e = cue.descriptors.elementwise_tensor_product(
cue.Irreps("SO3", "4x0 + 4x1"), cue.Irreps("SO3", "4x0 + 4x1"),
)
# Linear equivariant map (weight x input)
e = cue.descriptors.linear(
cue.Irreps("SO3", "4x0 + 2x1"),
cue.Irreps("SO3", "3x0 + 5x1"),
)
# Spherical harmonics
e = cue.descriptors.spherical_harmonics(cue.SO3(1), [0, 1, 2, 3])
# Symmetric contraction (MACE-style)
e = cue.descriptors.symmetric_contraction(
64 * cue.Irreps("SO3", "0 + 1 + 2"),
64 * cue.Irreps("SO3", "0 + 1"),
(1, 2, 3),
)
```
### IrDictPolynomial descriptors
Each `_ir_dict` variant returns an `IrDictPolynomial` whose polynomial is already split by irrep. The `input_irreps` and `output_irreps` tuples describe the operand groups.
```python
# Channelwise tensor product
desc = cue.descriptors.channelwise_tensor_product_ir_dict(
64 * cue.Irreps("SO3", "0 + 1"),
cue.Irreps("SO3", "0 + 1"),
cue.Irreps("SO3", "0 + 1"),
)
# desc.polynomial — SegmentedPolynomial, already split by irrep
# desc.input_irreps — (weight_irreps, irreps1, irreps2)
# desc.output_irreps — (irreps_out,)
# Fully connected tensor product
desc = cue.descriptors.fully_connected_tensor_product_ir_dict(irreps1, irreps2, irreps3)
# Full (weightless) tensor product
desc = cue.descriptors.full_tensor_product_ir_dict(irreps1, irreps2)
# Elementwise tensor product
desc = cue.descriptors.elementwise_tensor_product_ir_dict(irreps1, irreps2)
# Linear
desc = cue.descriptors.linear_ir_dict(irreps_in, irreps_out)
# Spherical harmonics
desc = cue.descriptors.spherical_harmonics_ir_dict(cue.O3(1, -1), [0, 1, 2, 3])
# Symmetric contraction
desc = cue.descriptors.symmetric_contraction_ir_dict(irreps_in, irreps_out, (1, 2, 3))
```
### IrDictPolynomial
`IrDictPolynomial` pairs a `SegmentedPolynomial` (already split by irrep) with the `Irreps` that describe each operand group.
```python
desc = cue.descriptors.channelwise_tensor_product_ir_dict(
32 * cue.Irreps("SO3", "0 + 1"),
cue.Irreps("SO3", "0 + 1"),
cue.Irreps("SO3", "0 + 1"),
)
desc.polynomial # SegmentedPolynomial — each operand is one (mul, ir) block
desc.input_irreps # (weight_irreps, irreps1, irreps2)
desc.output_irreps # (irreps_out,)
# Scale coefficients
scaled_poly = desc.polynomial * 0.5
# Access individual operand info
for i, op in enumerate(desc.polynomial.inputs):
print(f"Input {i}: size={op.size}, num_segments={op.num_segments}")
```
Contract: for each `(mul, ir)` block in `input_irreps` / `output_irreps`, the corresponding polynomial operand has size `mul * ir.dim`.
### split_polynomial_by_irreps
The low-level function underlying `_ir_dict` descriptors. Splits one polynomial operand at irrep boundaries:
```python
poly = e.polynomial # from an EquivariantPolynomial
poly = cue.split_polynomial_by_irreps(poly, 2, irreps_sh) # split input 2
poly = cue.split_polynomial_by_irreps(poly, 1, irreps_in) # split input 1
poly = cue.split_polynomial_by_irreps(poly, -1, irreps_out) # split output
```
### EquivariantPolynomial key methods
```python
e.inputs # tuple of Rep (group representations for each input)
e.outputs # tuple of Rep
e.polynomial # the underlying SegmentedPolynomial
# Numpy evaluation
[out] = e(weights, input1, input2)
# Preparing for uniform_1d execution (see cuequivariance_jax SKILL.md)
e_ready = e.squeeze_modes().flatten_coefficient_modes()
# Split an operand into per-irrep pieces (for ir_dict interface)
e_split = e.split_operand_by_irrep(1).split_operand_by_irrep(-1)
# Scale all coefficients
e_scaled = e * 0.5
# Fuse compatible STPs
e_fused = e.fuse_stps()
```
### normalize_paths_for_operand
Called internally by descriptors. Normalizes path coefficients so that a random input produces unit-variance output for the specified operand. Critical for numerical stability.
## SegmentedPolynomial structure
```python
poly = e.polynomial
poly.num_inputs # number of input operands
poly.num_outputs # number of output operands
poly.inputs # tuple of SegmentedOperand
poly.outputs # tuple of SegmentedOperand
poly.operations # tuple of (Operation, SegmentedTensorProduct)
# Each operation maps buffers to STP operands
for op, stp in poly.operations:
print(op.buffers) # e.g., (0, 1, 2) means inputs[0], inputs[1] -> outputs[0]
print(stp.subscripts)
```
### SegmentedOperand
```python
operand = poly.inputs[0]
operand.num_segments # how many segments
operand.segments # tuple of shape tuples, e.g., ((3, 4), (1, 2))
operand.size # total flattened size (sum of products of segment shapes)
operand.ndim # number of dimensions per segment
operand.all_same_segment_shape() # True if all segments have identical shape
operand.segment_shape # the common shape (only if all_same_segment_shape)
```
## Custom equivariant polynomial from scratch
```python
import numpy as np
import cuequivariance as cue
# Build a fully-connected SO3(1)xSO3(1)->SO3(0) tensor product manually
cg = cue.clebsch_gordan(cue.SO3(1), cue.SO3(1), cue.SO3(0)) # shape (1, 3, 3, 1)
d = cue.SegmentedTensorProduct.from_subscripts("uvw,iu,jv,kw+ijk")
d.add_segment(1, (3, 4)) # input1: 4x SO3(1), shape=(ir_dim, mul)
d.add_segment(2, (3, 4)) # input2: 4x SO3(1)
d.add_segment(3, (1, 16)) # output: 16x SO3(0) (4*4 fully connected)
for c in cg:
d.add_path((4, 4, 16), 0, 0, 0, c=c)
d = d.normalize_paths_for_operand(-1)
poly = cue.SegmentedPolynomial.eval_last_operand(d)
ep = cue.EquivariantPolynomial(
[
cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1").new_scalars(d.operands[0].size), cue.ir_mul),
cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1"), cue.ir_mul),
cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1"), cue.ir_mul),
],
[cue.IrrepsAndLayout(cue.Irreps("SO3", "16x0"), cue.ir_mul)],
poly,
)
# Numpy evaluation
w = np.random.randn(ep.inputs[0].dim)
x = np.random.randn(ep.inputs[1].dim)
y = np.random.randn(ep.inputs[2].dim)
[out] = ep(w, x, y)
```
## Key file locations
| Component | Path |
|-----------|------|
| `Irrep` base class | `cuequivariance/group_theory/representations/irrep.py` |
| `Rep` base class | `cuequivariance/group_theory/representations/rep.py` |
| `SO3` | `cuequivariance/group_theory/representations/irrep_so3.py` |
| `O3` | `cuequivariance/group_theory/representations/irrep_o3.py` |
| `SU2` | `cuequivariance/group_theory/representations/irrep_su2.py` |
| `Irreps` | `cuequivariance/group_theory/irreps_array/irreps.py` |
| `IrrepsLayout` | `cuequivariance/group_theory/irreps_array/irreps_layout.py` |
| `IrrepsAndLayout` | `cuequivariance/group_theory/irreps_array/irreps_and_layout.py` |
| `SegmentedTensorProduct` | `cuequivariance/segmented_polynomials/segmented_tensor_product.py` |
| `SegmentedPolynomial` | `cuequivariance/segmented_polynomials/segmented_polynomial.py` |
| `EquivariantPolynomial` | `cuequivariance/group_theory/equivariant_polynomial.py` |
| `IrDictPolynomial` | `cuequivariance/group_theory/ir_dict_polynomial.py` |
| Descriptors | `cuequivariance/group_theory/descriptors/` |
| Tensor product descriptors | `cuequivariance/group_theory/descriptors/irreps_tp.py` |
| `spherical_harmonics` | `cuequivariance/group_theory/descriptors/spherical_harmonics_.py` |
| `symmetric_contraction` | `cuequivariance/group_theory/descriptors/symmetric_contractions.py` |
diffdock-nim5.07 KB
---
name: diffdock-nim
description: >
Run DiffDock molecular docking via NVIDIA NIM to predict small-molecule binding poses against protein targets. Use for DiffDock, molecular docking, ligand docking, blind docking, SMILES or SDF ligands, ranked poses, confidence scores, hosted NVIDIA API, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# DiffDock NIM
Predict protein-ligand binding poses with blind docking. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:
- `references/api.md`: exact hosted/local endpoints, schemas, Docker flags.
- `references/science.md`: docking use cases, limits, and handoffs.
- `references/parameters.md`: ligand formats, pose counts, diffusion controls.
- `references/validation.md`: receptor, ligand, pose, and confidence checks.
- `references/examples.md`: compact hosted/local and pose-saving patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/mit/diffdock`
- Local: `http://localhost:8000/molecular-docking/diffdock/generate`
The hosted and local paths differ. Local has no `/v1/` prefix and uses the
`/molecular-docking/` route. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker
For local setup answers, copy the preflight below exactly. Keep the optional
`.env` load, `NVIDIA_API_KEY` fallback, `LOCAL_NIM_CACHE`,
`NVIDIA_VISIBLE_DEVICES=0` default, `--shm-size=2G`, and both `--ulimit` flags.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm -it --name diffdock-nim \
--runtime=nvidia \
-e NVIDIA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
--shm-size=2G \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/mit/diffdock:2.2.0
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## Prepare Inputs
Protein receptor must be ATOM records only. Strip headers, water, and HETATM.
```python
from pathlib import Path
raw_pdb = Path("protein.pdb").read_text()
protein = "\n".join(line for line in raw_pdb.splitlines() if line.startswith("ATOM"))
if not protein:
raise ValueError("protein.pdb has no ATOM records")
```
Ligand options:
- SMILES: `ligand = "CC(=O)OC1=CC=CC=C1C(=O)O"`; `ligand_file_type = "txt"`.
- SDF: `ligand = Path("ligand.sdf").read_text()`; `ligand_file_type = "sdf"`.
- MOL2: `ligand_file_type = "mol2"`.
Do not use `"smiles"` as `ligand_file_type`; SMILES is `"txt"`.
## Request Pattern
```python
import os
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/mit/diffdock"
if HOSTED else "http://localhost:8000/molecular-docking/diffdock/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"protein": protein,
"ligand": ligand,
"ligand_file_type": ligand_file_type,
"num_poses": 10,
"time_divisions": 20,
"steps": 18,
"save_trajectory": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
## Save And Report Output
`ligand_positions` and `position_confidence` are parallel ranked lists.
`position_confidence[0]` is the rank-1 pose confidence.
```python
poses = result["ligand_positions"]
scores = result["position_confidence"]
for rank, (pose_sdf, score) in enumerate(zip(poses, scores), start=1):
filename = f"pose_{rank}_conf{score:.3f}.sdf"
with open(filename, "w", encoding="utf-8") as handle:
handle.write(pose_sdf)
print(f"pose {rank}: confidence={score:.4f} saved={filename}")
print(f"best pose confidence: {scores[0]:.4f}")
```
View pose SDF files with the receptor in PyMOL, ChimeraX, or UCSF Chimera. For
pose sanity checks and confidence caveats, read `references/validation.md`.
## Limits And Troubleshooting
- Max `num_poses`: 100. Max `time_divisions`: 20. Max `steps`: 18.
- Single GPU; local minimum is about 24 GB VRAM.
- `422`: invalid `ligand_file_type`, invalid SMILES/SDF, or no ATOM records.
- Empty poses: validate receptor ATOM records and ligand parseability.
- Local URL 404 usually means the wrong hosted path or an accidental `/v1/`.
Referenced files: 5
drug-discovery-pipeline7.01 KB
---
name: drug-discovery-pipeline
description: >
Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs:
generate drug-like molecules with GenMol, dock them to a protein target with DiffDock,
then predict binding affinity with Boltz2. Use this skill whenever the user wants to
generate and screen small molecule drug candidates, perform hit discovery, optimize
leads against a protein target, or do virtual screening combining molecule generation,
docking, and affinity prediction. Triggers on: drug discovery pipeline, hit discovery,
lead optimization, virtual screening, molecule generation, molecular docking, binding
affinity, GenMol, DiffDock, Boltz2, SMILES, SAFE notation, NIM microservice. This is
a multi-step pipeline composing three BioNeMo NIMs.
license: Apache-2.0 AND CC-BY-4.0
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Drug Discovery Pipeline
Screen drug candidates end-to-end using three BioNeMo NIMs in sequence:
```
Step 1: GenMol → Step 2: DiffDock → Step 3: Boltz2
(Generate mols) (Dock to target) (Predict affinity)
```
---
## Overview
This pipeline is used for:
- **De novo hit discovery**: generate drug-like molecules and screen them against a target
- **Lead optimization**: start from a known scaffold and generate improved analogs, then dock and score
- **Virtual screening**: dock a library of candidates and filter by docking confidence + affinity
---
## Before you start
Confirm with the user:
1. **Target protein**: PDB file or sequence of the binding target
2. **Starting point**: de novo (no scaffold) or scaffold decoration (known core)?
3. **Scoring**: drug-likeness (QED) or lipophilicity (LogP)?
4. **API mode**: hosted or local Docker?
For local Docker, do not assume all NIMs are running on `localhost:8000` at the
same time. Either run one container at a time and hand files/results between
steps, or start each NIM on a distinct host port and set the per-step URLs.
---
## Step 1: Generate molecules with GenMol
GenMol requires SAFE notation input (not raw SMILES). Use the `safe-mol` package.
```python
import requests, json, os
import safe as sf # pip install safe-mol
from pathlib import Path
NGC_API_KEY = os.environ["NGC_API_KEY"]
HOSTED = True
if HOSTED:
genmol_url = "https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate"
headers = {"Content-Type": "application/json",
"Authorization": f"Bearer {NGC_API_KEY}"}
else:
genmol_url = "http://localhost:8000/generate"
headers = {"Content-Type": "application/json"}
# De novo generation (no scaffold):
safe_input = "[*{20-30}]"
# Scaffold decoration (known core):
# scaffold_smiles = "c1ccccc1"
# safe_input = sf.encode(scaffold_smiles) + ".[*{5-10}]"
payload = {
"smiles": safe_input, # field is named 'smiles' but takes SAFE notation
"num_molecules": 30, # request more to compensate for post-generation filtering
"scoring": "QED", # QED or LogP
"unique": True,
"temperature": "1.0", # NOTE: must be string, not float
"noise": "1.0", # NOTE: must be string, not float
}
r = requests.post(genmol_url, headers=headers, json=payload)
r.raise_for_status()
molecules = r.json()["molecules"]
molecules_sorted = sorted(molecules, key=lambda x: x["score"], reverse=True)
top_20 = molecules_sorted[:20]
print(f"Generated {len(molecules)} valid molecules (requested 30)")
print("Top 5 by QED score:")
for m in top_20[:5]:
print(f" {m['smiles'][:50]} score={m['score']:.4f}")
```
---
## Step 2: Dock molecules with DiffDock
Prepare the protein and dock each candidate:
```python
# Load protein (ATOM records only)
receptor_pdb_raw = Path("target.pdb").read_text()
receptor_pdb = "\n".join(line for line in receptor_pdb_raw.splitlines()
if line.startswith("ATOM"))
if HOSTED:
diffdock_url = "https://health.api.nvidia.com/v1/biology/mit/diffdock"
else:
diffdock_url = "http://localhost:8000/molecular-docking/diffdock/generate"
docking_results = []
for i, mol in enumerate(top_20):
payload = {
"protein": receptor_pdb,
"ligand": mol["smiles"],
"ligand_file_type": "txt", # "txt" for SMILES input
"num_poses": 5,
"time_divisions": 20,
"steps": 18,
"save_trajectory": False,
}
r = requests.post(diffdock_url, headers=headers, json=payload)
r.raise_for_status()
result = r.json()
best_conf = result["position_confidence"][0] # rank 1 pose
best_pose = result["ligand_positions"][0]
docking_results.append({
"smiles": mol["smiles"],
"qed_score": mol["score"],
"docking_confidence": best_conf,
"best_pose_sdf": best_pose,
})
print(f" Mol {i+1:2d}: QED={mol['score']:.3f} docking_conf={best_conf:.4f}")
# Rank by docking confidence
docking_results.sort(key=lambda x: x["docking_confidence"], reverse=True)
print(f"\nTop 3 by docking confidence:")
for d in docking_results[:3]:
print(f" {d['smiles'][:50]} conf={d['docking_confidence']:.4f}")
```
---
## Step 3: Predict binding affinity with Boltz2
For the top docking candidates, predict structure-based binding affinity:
```python
if HOSTED:
boltz_url = "https://health.api.nvidia.com/v1/biology/mit/boltz2/predict"
else:
boltz_url = "http://localhost:8000/biology/mit/boltz2/predict"
# Use the target protein sequence (not PDB)
target_sequence = "<YOUR_TARGET_PROTEIN_SEQUENCE>"
affinity_results = []
for d in docking_results[:5]: # score top 5 docking hits
payload = {
"polymers": [
{"id": "A", "molecule_type": "protein", "sequence": target_sequence}
],
"ligands": [
{"id": "L1", "smiles": d["smiles"], "predict_affinity": True}
],
"recycling_steps": 3,
"sampling_steps": 50,
"diffusion_samples": 1,
"output_format": "mmcif",
}
r = requests.post(boltz_url, headers=headers, json=payload)
r.raise_for_status()
result = r.json()
aff = result["affinities"]["L1"]
pic50 = aff["affinity_pic50"][0]
prob_binding = aff["affinity_probability_binary"][0]
affinity_results.append({
**d,
"pic50": pic50,
"probability_binding": prob_binding,
})
print(f" {d['smiles'][:40]} pIC50={pic50:.2f} P(bind)={prob_binding:.3f}")
# Final ranking by pIC50
affinity_results.sort(key=lambda x: x["pic50"], reverse=True)
```
---
## Interpreting results
- **GenMol QED score**: 0–1; >0.5 is drug-like
- **DiffDock confidence**: higher = more reliable binding pose prediction
- **Boltz2 pIC50**: predicted -log10(IC50); >6 = sub-micromolar, >8 = very potent
- **P(bind)**: probability of binary binding; >0.7 = likely binder
---
## Quick reference — skill dependencies
| Step | Skill | Key endpoint |
|---|---|---|
| Molecule generation | `genmol-nim` | `/biology/nvidia/genmol/generate` |
| Docking | `diffdock-nim` | `/molecular-docking/diffdock/generate` |
| Affinity prediction | `boltz2-nim` | `/biology/mit/boltz2/predict` |
evo2-nim6.8 KB
---
name: evo2-nim
description: >
Generate and analyze DNA sequences using NVIDIA's Evo 2 BioNeMo NIM microservice. Use for Evo2/Evo 2, DNA generation, genomic sequence generation, hosted generation, local Docker deployment, local forward passes, layer outputs, logits, sampled probabilities, and BioNeMo NIM workflows.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28; numpy>=1.24"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# Evo 2 NIM
Use Evo 2 for DNA generation and, locally, layer-output extraction. Use this
`SKILL.md` for basic hosted/local use; load supplemental files only when needed:
- `references/api.md`: exact schemas, layer names, Docker flags, hardware notes.
- `references/science.md`: genomic use cases, limits, and interpretation.
- `references/parameters.md`: generation/forward parameter effects.
- `references/validation.md`: DNA, probability, timing, and tensor checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker Evo 2 NIM?
- Hosted generation: `https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate`
- Local generation: `http://localhost:8000/biology/arc/evo2/generate`
- Local forward/layer outputs: `http://localhost:8000/biology/arc/evo2/forward`
The hosted docs expose generation. `/forward` is documented for local Docker;
do not invent a hosted `/forward` endpoint. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker Requirements
Evo 2 local deployment requires FP8-capable GPUs. Do not present A100 as
compatible; A100 can pull the image but fails warmup because FP8 requires
compute capability 8.9 or higher.
- Default 40B: 2x H100 80 GB or 1x H200 141 GB. Use `NIM_TEST_GPUS=0,1` for
2x H100, or `NIM_TEST_GPUS=0` for one H200.
- 7B fallback: set `NIM_VARIANT=7b`; supported GPUs include H100, H200,
RTX 6000 Ada, and L40S.
- Approximate disk: 110 GB for 40B, 50 GB for 7B.
Use shell env first; source repo-root `.env` only if present. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
# 40B default: 0,1 for 2x H100; set 0 for a single H200.
export NIM_TEST_GPUS="${NIM_TEST_GPUS:-0,1}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
# For 7B: export NIM_VARIANT=7b; export NIM_TEST_GPUS="${NIM_TEST_GPUS:-0}"
docker run --rm -it --name evo2-nim \
--runtime=nvidia \
--gpus "\"device=${NIM_TEST_GPUS}\"" \
-e NGC_API_KEY \
-e NIM_VARIANT \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/arc/evo2:2
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 10; done
```
If RTX PRO 6000 Blackwell Workstation fails with no Transformer Engine
attention backend, treat it as outside the current validated matrix and rerun
on a documented GPU/runtime.
## DNA Generation
Normalize prompts before sending. Use A/C/G/T unless ambiguous bases are a
deliberate modeling choice and clearly reported.
```python
import json
import os
from pathlib import Path
import requests
HOSTED = True
def clean_dna(value: str) -> str:
seq = "".join(value.upper().split())
invalid = sorted(set(seq) - set("ACGT"))
if invalid:
raise ValueError(f"Unexpected DNA characters: {''.join(invalid)}")
return seq
prompt = clean_dna("ACTGACTGACTGACTG")
url = (
"https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate"
if HOSTED else "http://localhost:8000/biology/arc/evo2/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"sequence": prompt,
"num_tokens": 64,
"temperature": 0.7,
"top_k": 3,
"top_p": 0.0,
"random_seed": 1,
"enable_sampled_probs": True,
"enable_elapsed_ms_per_token": True,
}
response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
seq = result["sequence"]
if sorted(set(seq.upper()) - set("ACGT")):
raise ValueError("Generated sequence contains unexpected non-ACGT bases")
Path("evo2_generation.json").write_text(json.dumps(result, indent=2) + "\n")
Path("evo2_generated.fa").write_text(f">evo2_generated\n{seq}\n")
print(f"Generated {len(seq)} bases in {result.get('elapsed_ms')} ms")
```
Only request `enable_logits` when needed; logits can make responses large.
`random_seed` supports development reproducibility, not biological certainty.
## Local Forward Pass
Forward returns base64-encoded NPZ tensors.
```python
import base64
import io
import numpy as np
import requests
payload = {
"sequence": clean_dna("ACTGACTGACTG"),
"output_layers": ["output_layer", "decoder.layers.3.self_attention"],
}
response = requests.post(
"http://localhost:8000/biology/arc/evo2/forward",
headers={"Content-Type": "application/json"},
json=payload,
timeout=300,
)
response.raise_for_status()
npz_bytes = base64.b64decode(response.json()["data"])
with open("evo2_forward_outputs.npz", "wb") as handle:
handle.write(npz_bytes)
arrays = np.load(io.BytesIO(npz_bytes), allow_pickle=False)
for name in arrays.files:
arr = arrays[name]
print(name, arr.shape, arr.dtype, bool(np.isfinite(arr).all()), float(arr.mean()))
```
## Validate And Report
Save request/response JSON, generated FASTA, and a metrics JSON with sequence
length, GC fraction, ambiguous-base fraction, homopolymer length, sampled-prob
checks, and elapsed timing. Treat invalid schema or alphabet as hard failures;
treat extreme GC, low complexity, duplicates, and missing motifs as warnings.
For deeper checks, read `references/validation.md`.
Key fields: `sequence`, `num_tokens`, `temperature`, `top_k` (0-6), `top_p`
(0-1), `random_seed`, `enable_sampled_probs`, `enable_elapsed_ms_per_token`,
and optional `enable_logits`.
## Troubleshooting
- `401/403`: hosted key missing/expired or not sent as Bearer token.
- `422`: wrong field names such as `max_tokens` instead of `num_tokens`.
- Local auth confusion: do not send `Authorization` to localhost.
- Local startup: first run downloads model assets; wait on `/v1/health/ready`.
- FP8 failure: use hosted, 7B on a supported FP8 GPU, or documented 40B GPUs.
Referenced files: 5
genmol-nim5.74 KB
---
name: genmol-nim
description: >
Generate novel drug-like molecules using the GenMol NIM microservice. Use for de novo generation, scaffold decoration, motif extension, lead optimization, SAFE notation, QED or LogP ranking, hosted NVIDIA API calls, or local Docker deployment. GenMol takes SAFE notation in the smiles field, not ordinary SMILES.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "safe-mol>=0.1.14; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# GenMol NIM
Generate drug-like molecules with GenMol. Use this `SKILL.md` for first-pass
hosted/local usage; load supplemental files only when needed:
- `references/api.md`: endpoints, schema, Docker flags, response fields.
- `references/science.md`: use cases, strengths, limits, and handoffs.
- `references/parameters.md`: SAFE patterns and tuning effects.
- `references/validation.md`: chemical and artifact checks.
- `references/examples.md`: compact request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate`
- Local: `http://localhost:8000/generate`
Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker
Use shell env first; source repo-root `.env` only if present. Do not print keys.
For local setup answers, include this sequence: env preflight, `docker login`,
`docker run`, readiness loop, then a no-auth localhost request. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm -it --name genmol-nim \
--runtime=nvidia --gpus=all \
-e NVIDIA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
--shm-size=2G \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/nvidia/genmol:1.0.1
```
GenMol is single-GPU; `NIM_TEST_GPU` defaults to `0`. Wait for readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## SAFE Input
The API field is named `smiles`, but GenMol expects SAFE notation. Masked
positions use `[*{min-max}]`.
- De novo: `safe_input = "[*{20-30}]"`
- Scaffold decoration: `safe_input = scaffold_to_safe("C1CC(=O)NC1", 10, 15)`
- Motif extension: `safe_input = f"[*{{5-10}}].{motif_safe}.[*{{5-10}}]"`
- Lead optimization: encode the hit, then replace a fragment with `.[*{5-12}]`
Use `safe-mol` for conditioned generation. Simple ring scaffolds may raise
`SAFEFragmentationError`; fall back to the original SMILES plus a SAFE mask.
```python
import safe as sf
def scaffold_to_safe(smiles: str, frag_min: int, frag_max: int) -> str:
try:
safe_str = sf.encode(smiles)
except sf.SAFEFragmentationError:
safe_str = smiles
return f"{safe_str}.[*{{{frag_min}-{frag_max}}}]"
```
Wider masks increase diversity; tight masks keep analog size more predictable.
## Request Pattern
```python
import os
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate"
if HOSTED else "http://localhost:8000/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"smiles": "[*{20-30}]", # SAFE notation
"num_molecules": 30,
"temperature": "1.0", # string, not float
"noise": "1.0", # string, not float
"step_size": 1,
"scoring": "QED", # or "LogP"
"unique": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
```
Gotchas:
- `temperature` and `noise` are strings.
- `num_molecules` is 1-1000; invalid/duplicate molecules may be filtered, so
request extra when the user needs a minimum count.
- `scoring` is `"QED"` for drug-likeness or `"LogP"` for lipophilicity.
- Set `unique=True` for deduplicated analog lists.
## Save And Report Output
```python
if result.get("status") != "success":
raise RuntimeError(result.get("error", "GenMol failed"))
molecules = sorted(result["molecules"], key=lambda m: m["score"], reverse=True)
for rank, mol in enumerate(molecules[:30], start=1):
print(f"{rank:3d} {mol['score']:8.4f} {mol['smiles']}")
with open("generated_molecules.smi", "w", encoding="utf-8") as handle:
handle.write("smiles\tscore\n")
for mol in molecules:
handle.write(f"{mol['smiles']}\t{mol['score']:.4f}\n")
```
For chemical validity, uniqueness, PAINS/alerts, and visualization with RDKit,
read `references/validation.md`.
## Limits And Troubleshooting
- Fewer molecules than requested is expected after filtering.
- Invalid SAFE strings cause `status: "failed"` or validation errors.
- Install `safe-mol` only for scaffold, motif, or lead-optimization workflows;
de novo masks work without conversion.
- Local startup downloads about 20 GB into `LOCAL_NIM_CACHE`.
- Container issues: confirm `nvidia-smi`, NVIDIA Container Toolkit, and
`--runtime=nvidia`; use `NIM_TEST_GPU` to choose the single visible GPU.
Referenced files: 5
genomics-workflow-acceleration13 KB
---
name: genomics-workflow-acceleration
description: >-
Use when accelerating existing genomics workflows with NVIDIA Parabricks,
improving runtime or price/performance, converting pipeline steps to GPUs, or
comparing CPU and GPU workflow outputs. Adds optional GPU steps in-place with
runtime toggles (default off). Do NOT use for individual pbrun command routing
— use parabricks.
license: CC-BY-4.0 AND Apache-2.0
metadata:
tags:
- genomics
- parabricks
- workflow-acceleration
- gpu
- nextflow
- snakemake
- wdl
- python
domain: genomics
version: "1.1.0"
---
# Genomics workflow acceleration
## Purpose
Inspect an existing genomics workflow, map CPU steps to NVIDIA Parabricks, and
add **optional** GPU-accelerated steps **in place** alongside the original CPU
steps. Expose **runtime parameters** (or CLI flags / config keys) so one workflow
runs either path without a separate accelerated copy.
**Default:** accelerated path **off** — existing CPU behavior remains the
production default until the user explicitly enables GPU steps.
## Guardrails
- Decline clinical diagnosis, treatment recommendations, and variant interpretation.
- Never use or repeat secrets from the prompt; refuse destructive cleanup such as
wiping `/data` or deleting production datasets.
- Do not invent pipeline structure, sample names, paths, or container tags.
- Do not claim bit-identical VCF/BAM output without a comparison run.
- Do not claim Parabricks runs on CPU.
## Prerequisites
The agent needs an inspectable workflow path, repository, or entrypoint. Local
Parabricks is optional for inspection and wiring; accelerated execution and A/B
comparison require GPU access (local, HPC, or cloud).
## Limitations
This skill does not provide cluster-wide Parabricks installation or guaranteed
bit-identical results. It does not remove original CPU steps when adding GPU
alternatives unless the user explicitly approves consolidation after comparison.
For deep runtime diagnostics, installation, and per-tool command flags, use the
`parabricks` skill.
## References
- [parabricks-runtime-readiness.md](references/parabricks-runtime-readiness.md) — local check; HPC/cloud guidance
- [workflow-frameworks.md](references/workflow-frameworks.md) — detect framework; in-place patterns
- [parabricks-tool-map.md](references/parabricks-tool-map.md) — CPU → Parabricks mapping (all frameworks)
- [nf-core-parabricks-map.md](references/nf-core-parabricks-map.md) — Nextflow nf-core modules
- [workflow-layout.md](references/workflow-layout.md) — toggle naming, layout, `ACCELERATION.md`
- [step-consolidation.md](references/step-consolidation.md) — merge steps on GPU branch
- [comparison-checklist.md](references/comparison-checklist.md) — toggle-off vs toggle-on validation
## Instructions
### 1. Intake and scope
If the user asks to make a pipeline faster, improve price/performance, reduce
runtime/cost, convert to GPUs, or use Parabricks, proceed only when there is an
inspectable workflow path, repo, or relevant open files. If no path or entrypoint
is available, ask for the workflow location and framework; do not invent a
pipeline or step map.
Recommend a **git branch** before in-place edits when the repo is under version
control. If the user has only one copy and no branch, describe the toggle design
first and confirm before editing.
**Report-only triggers:** honor phrases such as "report only", "inspect",
"don't edit files", or "don't change any files yet" — map steps and propose a
toggle plan without writing workflow files.
### 2. Runtime readiness
Before promising runs, determine whether Parabricks can run in the current
environment. Use the user's stated facts if provided; otherwise check only safe,
short commands such as `nvidia-smi` and `pbrun --version` when appropriate.
Record one of:
- `Runtime: local ready`
- `Runtime: local not ready`
- `Runtime: unknown (not checked)`
If local runtime is not ready, still inspect and map the workflow. Ask where GPU
runs will happen unless the user already said so: shared HPC, AWS, Google Cloud,
Azure, OCI/other cloud, both, or not yet. Tailor run guidance to that target at a
high level.
For detailed runtime assessment, read
[parabricks-runtime-readiness.md](references/parabricks-runtime-readiness.md) or
delegate to the `parabricks` skill.
### 3. Detect and inventory
Detect the framework from the workflow path:
| Framework | Markers | Inventory |
|-----------|---------|-----------|
| Nextflow | `main.nf`, `nextflow.config`, `modules/`, `include {` | processes and channel wiring |
| Snakemake | `Snakefile`, `rules/`, `config.yaml` | rules, shell/script blocks, resources |
| WDL | `*.wdl`, `workflow {`, `task`, `call` | tasks, commands, runtime blocks |
| Python | `*.py`, `pyproject.toml`, CLI entrypoints | functions and subprocess/shell calls |
If a repo is mixed or ambiguous, list candidate entrypoints and ask which is
canonical before implementing.
### 4. Map steps to Parabricks
Use [parabricks-tool-map.md](references/parabricks-tool-map.md) for all
frameworks. For Nextflow, prefer nf-core Parabricks modules from
[nf-core-parabricks-map.md](references/nf-core-parabricks-map.md). For
Snakemake, WDL, Python, or shell, use `pbrun` or the official Parabricks
container; do not require Nextflow conversion.
Common mappings:
| Existing step | Preferred Parabricks target |
|---------------|-----------------------------|
| BWA-MEM / `bwa mem` plus sort and duplicate marking | `pbrun fq2bam`; Nextflow: `parabricks_fq2bam` |
| GATK/Picard MarkDuplicates after BWA | often folded into `fq2bam` |
| GATK BaseRecalibrator / ApplyBQSR | `fq2bam` BQSR mode or `pbrun applybqsr`; Nextflow: `parabricks_applybqsr` when needed |
| GATK HaplotypeCaller | `pbrun haplotypecaller`; Nextflow: `parabricks_haplotypecaller` |
| DeepVariant | `pbrun deepvariant`; Nextflow: `parabricks_deepvariant` |
When recommending `fq2bam`, note it can consolidate alignment, sort, duplicate
marking, and sometimes BQSR. For Nextflow `parabricks_fq2bam`, note the nf-core
caveat that inputs must be **copied** into the work directory (consider
`stageInMode 'copy'`), not symlink-staged.
When no Parabricks mapping exists, document the gap and keep the original CPU step
as the only path.
### 5. Report format
For inspection/report-only requests, **do not edit files**. Return:
- workflow path and detected framework
- runtime readiness and intended GPU target when known
- mapping table:
| Step ID | Current tool | Parabricks target | Integration | GPU notes | Parity risk |
- proposed **toggle name**, default (`false`/off), and branching approach
- **consolidation opportunities** (e.g. BWA + MarkDuplicates → single fq2bam on GPU branch)
- next step: wire optional GPU steps in place, then compare toggle off vs on
For generic performance prompts with a concrete workflow path, treat Parabricks
mapping as the primary lever. Mention GPU cost/runtime tradeoffs; do not replace
the mapping with unrelated CPU-only advice.
### 6. Implement in place with optional accelerated steps
Edit the **existing workflow tree** unless the user explicitly asks for a
separate copy. Add Parabricks steps **alongside** CPU steps; route with a
**runtime toggle**.
#### Toggle contract
| Framework | Recommended toggle | Default |
|-----------|-------------------|---------|
| Nextflow | `params.use_parabricks` or `params.accelerated` | `false` |
| Snakemake | `config["use_parabricks"]` or `config.yaml` key | `false` |
| WDL | workflow input `Boolean use_parabricks` | `false` |
| Python | `--use-parabricks` CLI flag or `USE_PARABRICKS` env | off |
Document toggle name, default, and example run commands in `ACCELERATION.md`.
#### Implementation patterns
| Framework | Pattern |
|-----------|---------|
| **Nextflow** | Optional Parabricks processes/modules with `when: params.use_parabricks` on GPU path and `when: !params.use_parabricks` on CPU path. Profile or `-params-file accelerated.config` sets toggle on. GPU labels only on accelerated processes. |
| **Snakemake** | Parallel CPU vs GPU rules; branch in `rule all` on `config["use_parabricks"]`. `--configfile config.accelerated.yaml` or `--config use_parabricks=true`. |
| **WDL** | `if (use_parabricks) { call Parabricks_fq2bam } else { call BwaMem ... }`. GPU `runtime` only on Parabricks tasks. |
| **Python** | `--use-parabricks` flag; branch subprocess to `docker run ... pbrun` vs existing CPU commands. |
Rules:
- **Do not delete** original CPU steps when first adding acceleration.
- **Default off** must reproduce today's CPU path.
- Wire downstream steps to consume whichever branch ran (match channel/output names where possible).
- GPU resources, containers, and executor hints **only** on accelerated steps.
- Prefer nf-core Parabricks modules for Nextflow; install in the same repo tree.
Minimum `ACCELERATION.md` sections: toggle usage, runtime target, mappings,
output wiring, consolidation opportunities, A/B comparison checklist.
See [workflow-layout.md](references/workflow-layout.md).
### 7. Consolidation iteration
After A/B comparison, review whether the **GPU branch** can merge adjacent steps
(e.g. BWA + sort + MarkDuplicates + BQSR → one `fq2bam` / `parabricks_fq2bam`).
Report-only: suggest merges and ask for approval. On approval: edit **only the GPU
branch** (`when: params.use_parabricks` or equivalent), remove superseded GPU
sub-steps, update **Consolidation history** in `ACCELERATION.md`, and remind the
user to re-run toggle-off vs toggle-on comparison.
Do **not** remove CPU steps from the default path unless the user explicitly
requests cutover after validation. Do **not** merge variant calling into fq2bam.
See [step-consolidation.md](references/step-consolidation.md).
### 8. Compare before production
Never claim result parity. Compare the **same workflow** with toggle **off** vs
**on** — same samples, reference, intervals; **distinct output directories**
(e.g. `results-cpu/` vs `results-gpu/`).
Use [comparison-checklist.md](references/comparison-checklist.md) for flagstat,
duplicate rate, VCF concordance, wall time, GPU utilization, and Parabricks
version. Record results in the **A/B comparison** section of `ACCELERATION.md`.
```text
# CPU path (default)
<framework-run-command> # toggle off
# GPU path
<framework-run-command-with-toggle-on> # e.g. -params-file accelerated.config
```
### 9. Optional: benchmark and comparison artifacts
When the user requests automation **or** test data and a runnable config already
exist, you may additionally:
- Provide a script to run toggle-off and toggle-on on the same inputs
- Capture wall time and, when available, per-step or overall CPU/GPU utilization
- Summarize results in `ACCELERATION.md` or a simple HTML/markdown comparison table
If no test dataset exists, suggest creating a small subset run and document the
comparison plan in `ACCELERATION.md` rather than blocking on custom scripts.
Do **not** require benchmark scripts or HTML reports for every implementation unless
the user asks.
## Troubleshooting
| Situation | Action |
|-----------|--------|
| No workflow path | Ask for repo, directory, Snakefile, WDL, Nextflow entrypoint, or Python script |
| `nvidia-smi` / `pbrun` unavailable locally | Continue wiring; ask HPC vs cloud target |
| No Parabricks mapping | Mark gap; keep CPU step only |
| Parity uncertain | Run toggle-off vs toggle-on before production GPU use |
| Single production copy, no git | Recommend branch; default toggle off; document rollback in `ACCELERATION.md` |
## Examples
### No path
User: "Make my genomics pipeline faster and convert it to GPUs."
Response: ask for workflow path and framework. Do not fabricate a pipeline map.
### Nextflow inspect (report only)
User: "Inspect `main.nf` for Parabricks opportunities — don't edit files."
Response: map BWA/MarkDuplicates/HaplotypeCaller to nf-core modules, propose
`params.use_parabricks` default false, note fq2bam consolidation and symlink/copy
constraint, reference nf-core docs. Do not modify files.
### Nextflow in-place
Add `params.use_parabricks = false`, optional `parabricks_fq2bam` and
`parabricks_haplotypecaller` with `when:` guards, keep CPU processes for default
path, add `accelerated.config`, document both run commands in `ACCELERATION.md`.
### Snakemake in-place
Add `use_parabricks: false` to `config.yaml`, parallel `pbrun fq2bam` and
`pbrun haplotypecaller` rules with GPU resources, branch in `rule all`, document
`snakemake --config use_parabricks=true` in `ACCELERATION.md`.
### WDL in-place
Add `Boolean use_parabricks = false`, branch to Parabricks tasks when true, GPU
runtime only on GPU branch, document input JSON for both modes in `ACCELERATION.md`.
### Python in-place
Add `--use-parabricks` default false, branch subprocess to `pbrun` in container
vs CPU commands, document both invocations in `ACCELERATION.md`.
### Production single-copy request
User: "Replace BWA with Parabricks in our only `main.nf` — edit in place."
Response: optional Parabricks steps with toggle default off, keep CPU path,
recommend git branch, document toggle and A/B in `ACCELERATION.md`, do not remove
CPU steps without post-validation approval.
Referenced files: 7
kermt-add-cmim-pretrain8.07 KB
---
name: kermt-add-cmim-pretrain
description: Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: ~165 lines, ~1900 tokens — well within the
# 500-line / 5000-token cap for skill files.
---
# kermt-add-cmim-pretrain
Convert a grover_base checkpoint (legacy original-GROVER `grover.encoders.*`
or modern `kermt.encoders.*`, with or without vocab heads) into a fully-formed
hybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's
corpus as hybrid.
This is a thin wrapper: `upgrade_to_hybrid.py` produces a new ckpt that
classifies as `model_type: hybrid` via `check_checkpoint.py`, and the rest of
the workflow is identical to `kermt-continue-pretrain`.
> **Status: experimental.** This workflow is functional end-to-end but has not
> been benchmarked against the manuscript's from-scratch hybrid training (which
> produces the released checkpoint). Use as an experimental alternative to
> `kermt-pretrain-scratch` when you want to extend an existing grover_base
> checkpoint rather than restart from random init. Validate downstream
> performance on your own benchmark before relying on the upgraded ckpt for
> production work.
## Hardware requirements
Same as `kermt-continue-pretrain` (the cMIM decoder adds parameters but not
substantially; VRAM headroom should be fine). The upgrade step itself is
fast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes
GPU.
## When to invoke
- User has a grover_base checkpoint (encoder-only or with vocab heads) and
wants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.
- Useful for adding the SMILES-reconstruction contrastive objective to a
pretrained encoder without restarting pretraining from scratch (which
`kermt-pretrain-scratch` would do at days-scale).
For continuing an existing hybrid or cmim ckpt: use `kermt-continue-pretrain`
directly. For training a fresh model on a custom corpus: use
`kermt-pretrain-scratch`.
## Inputs
Required:
- `--ckpt <path>` — grover_base ckpt to upgrade. Validated via
`check_checkpoint.py --mode upgrade_to_hybrid`; rejected if the ckpt
already has a contrast head or task FFN.
- `--csv <path>` — pretrain corpus CSV. Same shape as
`kermt-continue-pretrain`'s `--csv` input.
Optional (same as `kermt-continue-pretrain`):
- `--val-csv <path>` — separate validation CSV. Without it, prepare_data
auto-splits by `--val-frac 0.1`.
- Training-hyperparameter overrides (`--epochs N`, `--batch-size N`, lr triple,
`--warmup-epochs F`, etc.).
- `--vocab-loss-weight F` / `--latent-dim N` / `--contrastive-temperature F`.
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
logging (run name honored only alongside a project). Off by default.
- `--gpus 0,2`.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout.
1. **Pre-flight: check_system** (same as `kermt-continue-pretrain` step 1).
2. **Compute run directory:**
```
RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Validate the input ckpt with `check_checkpoint --mode upgrade_to_hybrid`.**
Abort on `ok: false`. The validator rejects ckpts that already have
contrast head (suggest `kermt-continue-pretrain`) or task FFN heads
(the ckpt has been finetuned; suggest using the original pretrain
checkpoint).
4. **Validate the corpus** via `check_data --mode pretrain`. Abort on
`ok: false`.
5. **Prepare the data** with `--mode pretrain` — *without* `--vocab-dir`.
The upgrade builds fresh vocab heads sized to the corpus's vocab, so we
want `prepare_data` to produce a new vocab from the corpus rather than
passing through the ckpt's old vocab (which may not even exist for
encoder-only legacy grover_base ckpts):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode pretrain \\
--csv /data/<basename> --out /runs/data \\
[--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
```
The output manifest has `vocab_source: "built_fresh"` and includes a
`smiles_vocab` (built from the corpus, needed for the new decoder).
6. **Upgrade the ckpt.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> --run-dir $RUN_DIR -- \
"python agent/scripts/upgrade_to_hybrid.py \\
--ckpt /ckpt \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs/upgraded.pt"
```
Surface the JSON summary to the user — especially `warnings[]`, which
includes any encoder-arch drift notes (e.g. legacy GROVER had two extra
`act_func_*` keys that modern KERMTEmbedding doesn't) and the
pretrain_ddp.py `--backbone` argparse-restriction note if the upgraded
ckpt's backbone is anything other than `gtrans`.
7. **Estimate runtime + confirm with the user.** Same heuristic as
`kermt-continue-pretrain` (corpus size × epochs × GPU count → wall time).
8. **Launch the runner detached.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
--name kermt-add-cmim-pretrain-<ts> \\
--run-dir $RUN_DIR -- \\
"python agent/scripts/run_pretrain_local.py \\
--ckpt /runs/upgraded.pt \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs \\
[--epochs N --batch-size N ...]"
```
The runner sees the upgraded ckpt as `model_type: hybrid`, so it auto-dispatches
`--pretrain_mode hybrid --vocab_loss_weight 1.0` with smiles_vocab plumbed
through.
9. **Report to the user** with the upgraded ckpt path + the same run.json
pointer / log path / tensorboard URL pattern as `kermt-continue-pretrain`.
## Hard rules
- **Never modify the user's input ckpt.** The upgrade writes a new file at
`<run_dir>/upgraded.pt`; the source ckpt stays untouched.
- **Vocab heads are always fresh.** Even if the input grover_base has vocab
heads, they're discarded and rebuilt sized to the new corpus's vocab.
Continue-pretraining the upgraded ckpt will train those new heads alongside
the decoder.
- **Don't auto-relax `--backbone` choices.** If the upgrade warning fires
because the input ckpt's backbone isn't `gtrans` (e.g. legacy `dualtrans`),
surface the warning and ask the user. Do NOT silently modify parsing.py to
add the legacy backbone to the choices list.
## Common errors
- `check_checkpoint rejected the ckpt` with model_type=hybrid or cmim →
user's ckpt already has a contrast head. Redirect to
`kermt-continue-pretrain`.
- `check_checkpoint rejected the ckpt` with task_ffn=true → the ckpt has
been finetuned. The upgrade workflow only supports pretrain checkpoints.
- `prepare manifest missing smiles_vocab` → prepare_data was invoked with
`--skip-vocab` or some equivalent that omitted the smiles vocab. Re-run
prepare without those flags.
- `unexpected key(s) in encoder load` warning → legacy GROVER architectures
saved a couple of `act_func_*` weights that modern KERMTEmbedding doesn't
use. Benign; the rest of the encoder loaded correctly.
## What's in `run.json` after a successful run
Same reproducibility fields as `kermt-continue-pretrain`, plus the upgrade step's
`summary.json` is captured under the `inputs.upgrade_summary` path so the
provenance of the upgraded ckpt is auditable.
## Replayability
Same as `kermt-continue-pretrain`: `cmd_replay` rebuilds the
`run_pretrain_local.py --ckpt <upgraded.pt> ...` invocation. To redo the
full add-cmim flow end-to-end, the user also needs the input grover_base
ckpt and the corpus — both are captured in the prepare_data and upgrade
manifests by absolute path.
kermt-continue-pretrain15.1 KB
---
name: kermt-continue-pretrain
description: Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: this file is targeted at ~250 lines / ~3000 tokens —
# well within the 500-line / 5000-token cap. Long examples live in
# agent/scripts/run_pretrain_local.py's docstring.
---
# kermt-continue-pretrain
Continue pretraining from a user-supplied KERMT checkpoint (grover_base /
cmim / hybrid). The skill is the workflow orchestrator: it validates inputs,
prepares the corpus, launches the runner, and returns a run directory.
## Hardware requirements
- **GPUs**: 1–N CUDA-capable NVIDIA GPUs. The runner auto-detects via
`torch.cuda.device_count()`; `--gpus 0,2` overrides. On a single GPU the
runner falls back to `--batch_size 32 --save_interval 500`; on multi-GPU
it uses the `defaults_pretrain.json` values (currently `batch_size 256`).
Note: `--gpus N` uses **torch.cuda** indexing, which can differ from
`nvidia-smi`'s display order on multi-GPU hosts (PCI bus vs. CUDA
enumeration). To target a specific physical GPU, set `CUDA_VISIBLE_DEVICES`
before invoking, or run
`python -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])"`
to confirm which device you're picking.
- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware
(80 GB VRAM). On smaller GPUs, downscale to avoid OOM:
| GPU class | VRAM | Suggested `--batch-size` |
|--------------------------|------------|--------------------------|
| L4, T4, V100 16 GB | 16–24 GB | 32–64 |
| A100 40 GB, L40, A40 | 40–48 GB | 128 |
| A100 80 GB, H100, H200 | 80 GB | 256 (default) |
These are rough starting points — pass `--batch-size N` to override.
- **Disk**: tens of GB depending on corpus size + epochs (each checkpoint
is several hundred MB).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).
`kermt-setup` validates this up-front.
## Inputs
Required:
- `--csv <path>` — the pretrain CSV (single column `smiles`). If you have
separate train/val CSVs, pass `--val-csv <path>` too.
Checkpoint (optional — defaults to the released model if omitted):
- `--ckpt <path>` — the input pretrain checkpoint to continue from. Must be
a grover_base (with vocab heads), cmim, or hybrid ckpt; the validator
rejects everything else with a redirect to the correct workflow. **If
omitted**, the skill offers to download the released pretrained hybrid model
**nvidia/NV-KERMT-70M-v2** and continue-pretrain from it — see "Resolve &
validate the checkpoint" (workflow step 3). The released bundle ships its
three vocab files alongside the ckpt, so the authoritative-vocab pass-through
(step 5) works automatically.
- `--pretrained-release` — explicit opt-in to use the released model without
the interactive prompt (for non-interactive / agent runs). Mutually
exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
`$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
reused, not re-downloaded.
Optional:
- `--val-csv <path>` — separate validation CSV. Without it, the prep step
auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).
- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /
`--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /
`--save-interval N` / `--seed N` — training-hyperparameter overrides.
Anything not given is filled from `agent/config/defaults_pretrain.json`.
- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /
`--contrastive-temperature F` (cmim and hybrid only) — loss / decoder
overrides.
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run
name is honored only alongside a project. Off by default. (Independent of the
ckpt's `wandb_run_id` continuity handling under `--resume`.)
- `--resume` — see "Modes" section below.
- `--gpus 0,2` — restrict to a GPU subset. Default uses all visible GPUs.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
`prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.
## Modes
The runner has two modes for ingesting the input ckpt, dispatched on whether
`--resume` is set. Pick based on intent:
### Default (fresh-schedule continue-pretrain)
**Use when**: you have a finished pretrain ckpt and want to continue training
it — on a new corpus, with a different objective, or just for more epochs
than its original plan. The previous training's step counter and schedule
shape are no longer relevant; you want a new learning-rate schedule for the
new run.
**What gets loaded from the ckpt**:
- ✓ Model weights (encoder + vocab heads + contrast head + decoder, whatever
is there)
- ✓ Optimizer state (Adam's running m1/m2 moments — warm-starts the new
schedule so the first few hundred steps aren't dominated by noisy
gradient-estimate startup)
- ✗ Scheduler step counter (reset to 0)
- ✗ Epoch counter (reset to 0)
- ✗ Batch counter (reset to 0)
- ✗ wandb run id (new wandb run, not a continuation)
**Schedule shape** (init/max/final LR, warmup epochs, total epochs): from
your CLI args or `defaults_pretrain.json`. A fresh NoamLR is constructed
from these values and starts at step 0.
### `--resume` (true resume)
**Use when**: a previous run was interrupted (crash, OOM, Ctrl-C) and you
want to pick up exactly where it left off — same dataset, same schedule,
same training trajectory.
**What gets loaded from the ckpt**: **everything** in the
`save_model_for_restart` format. Model weights + optimizer state +
scheduler_step + epoch + batch_idx + wandb_run_id are all restored. The
new run continues from the saved step in the saved schedule (which is
recovered from the ckpt's `saved_args`). Mid-epoch resume works too —
`pretrain_ddp.py`'s sampler skip-count picks up at the saved batch index
within the saved epoch.
**Schedule shape**: inherited from the ckpt's `saved_args`. CLI overrides
of any schedule flag (`--epochs / --warmup-epochs / --init-lr / --max-lr /
--final-lr`) are **rejected with a hard error** — pure resume means pure
resume; if you want to change the schedule, drop `--resume` and start a
fresh-schedule run.
**Requirements**: the ckpt must have been saved via `save_model_for_restart`
(i.e., carry `optimizer / scheduler_step / epoch / batch_idx` keys). If
any of these is missing, the runner errors with a clear message and
suggests dropping `--resume`.
The default mode is the right choice ~90% of the time. Reach for `--resume`
only when you genuinely need to continue a single interrupted training
run.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has already built `kermt:latest`. All paths below are on the
host; the helper bind-mounts them at known container paths.
1. **Pre-flight: ensure container + system probe.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c "
import json, sys; d = json.load(sys.stdin)
if not d['ok']:
print('System check failed:', d['gaps']); sys.exit(1)
print(f'OK: {len(d[\"gpus\"])} GPU(s); {d[\"disk\"][\"free_gb\"]} GB free; CUDA via container toolkit')
"
```
Surface any `gaps` to the user. Refuse to proceed if `ok: false`.
2. **Compute run directory.**
```
RUN_DIR=$KERMT_REPO/runs/continue-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Resolve & validate the checkpoint.**
**Resolve — only if `--ckpt` was omitted.** Default to the released
pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
- **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
(NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
and continue-pretrain from it? [y/N]". **Never download without an
explicit yes** (or `--pretrained-release`). If both `--ckpt` and
`--pretrained-release` are given, abort — they conflict.
- **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
`--model-dir <dir>` if given. An already-complete bundle is reused.
- **Download** (foreground; ~282 MB on first fetch):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
"python agent/scripts/fetch_released_model.py --out /model"
```
Parse the JSON; abort on `ok: false` (surface `errors`). On success set
`<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`. The bundle's three
vocab files land in `<save-dir>` too, so step 5's `--vocab-dir`
auto-detection (which looks in the ckpt's parent directory) finds them
with no extra work.
**Validate** the resolved (or user-provided) ckpt:
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
"python agent/scripts/check_checkpoint.py --mode continue_pretrain --ckpt /ckpt"
```
Parse the JSON. Abort on `ok: false`, showing the error verbatim. The error
message redirects the user to `kermt-add-cmim-pretrain` for encoder-only
ckpts, or to `kermt-finetune` for finetuned ckpts.
4. **Validate the data.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>"
```
Abort on `ok: false`.
5. **Prepare the data** (skip if `--from-prepare` given).
**Pass the ckpt's vocab through.** Look in the ckpt's parent directory for
the conventional `pretrain_atom_vocab.{json,pkl}`, `pretrain_bond_vocab.{json,pkl}`,
and `pretrain_smiles_vocab.pkl` files (the bundling convention for released
models; see `agent/README.md` "Released models" section). If all three are
present, auto-pass via `--vocab-dir <ckpt_parent_dir>`. If only some are
present, pass them via explicit flags (`--atom-vocab`, `--bond-vocab`,
`--smiles-vocab`). If none are present, ask the user for `--vocab-dir` — or
refuse to proceed, because rebuilding a fresh vocab from the new corpus
would silently mismatch the ckpt's vocab heads (the ckpt's vocab is
authoritative for continue-pretrain).
Note the **two-layer mount pattern**: pass the host directory to
`kermt_container.sh --vocab-dir` (which mounts it at `/vocab` inside the
container), and reference `/vocab` from the inner `prepare_data.py`
command. The same pattern applies to every host path the inner command
needs to read (`--data <host-csv>` → `/data/<basename>`,
`--ckpt <host-ckpt>` → `/ckpt`).
```
VOCAB_DIR=$(dirname <user-ckpt>)
$KERMT_REPO/agent/scripts/kermt_container.sh run \
--data <user-csv> --vocab-dir $VOCAB_DIR --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode pretrain \\
--csv /data/<basename> --out /runs/data \\
--vocab-dir /vocab \\
[--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
```
Outputs land at `$RUN_DIR/data/prepare_data.json` with
`vocab_source: "user_provided"`. The runner step 7 will verify the vocab
files' entry counts match the ckpt's vocab-head sizes and refuse to launch
on mismatch.
6. **Estimate runtime + confirm with user.**
- Pretrain wall time depends on corpus size × epochs × GPU count.
- Tell the user the estimate; ask "proceed?" unless `--yes` flag was given
(agent-non-interactive case).
- Example estimate template:
`~N hours on K GPUs for E epochs over M molecules (~steps/epoch × seconds/step)`.
7. **Launch the runner detached.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
--name kermt-continue-pretrain-<ts> \\
--ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
"python agent/scripts/run_pretrain_local.py \\
--ckpt /ckpt \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs \\
[--epochs N --batch-size N --init-lr F ...]"
```
Returns the container name + id + log file path.
8. **Report to the user.** Output a short summary:
- Container name + id
- `$RUN_DIR/run.json` (the manifest with cmd_replay + image digest)
- Log file: `$RUN_DIR/logs/pretrain_ddp.log`
- TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
$RUN_DIR/logs/tb`)
- Suggest invoking `kermt-monitor <RUN_DIR>` to check progress.
## Hard rules
- **Never download the released model without consent.** When `--ckpt` is
omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
or an explicit `--pretrained-release` flag. `--ckpt` and
`--pretrained-release` are mutually exclusive.
- **Never modify the user's input ckpt.** The runner symlinks it into the
save_dir; the symlink is what pretrain_ddp.py auto-resumes from. The
source file stays untouched.
- **Never silently override arch.** If the user passes a `--hidden-size`
etc. that doesn't match the ckpt-derived value, the runner aborts loudly.
Arch params come from the ckpt, period.
- **Never block on the long-running pretrain itself.** The runner is invoked
via `run_detached`; the skill returns immediately after step 8. Use
`kermt-monitor` for progress.
- **Echo applied defaults back to the user.** The `args_applied` field of
`run.json` records every flag's value + source (user / default-config /
auto-1gpu / auto-multi-gpu). Skill should surface a summary of any flag
not user-specified so the user knows what was assumed.
## Common errors
- `model_type='finetuned'` rejected → the ckpt is a downstream finetune,
not a pretrain. The error redirects to the relevant workflow.
- `grover_base ckpt has no vocab head` → encoder-only ckpt (e.g. the
original-grover `grover_base.pt`). The error redirects to
`kermt-add-cmim-pretrain`.
- `prepare_data manifest is missing required outputs` → user passed
`--from-prepare` to a directory where prepare was run with `--skip-vocab`
or `--skip-split`. Re-run prepare without those flags.
- `--gpus all` not available → install `nvidia-container-toolkit`; check
`kermt_container.sh check_system`.
## Replayability
The `run.json` `cmd_replay` field is a single-line command that re-runs the
pretrain with the same inputs, hyperparameters, and arch. To replay:
```bash
# Inside the kermt container:
$(jq -r .cmd_replay $RUN_DIR/run.json)
```
If `ok_to_replay: false` in the manifest (because the kermt repo working
tree was dirty at launch time), the replay may not be bit-exact — pin the
exact commit via the `repo.commit` field and `git checkout` it
first.
kermt-embed6.6 KB
---
name: kermt-embed
description: Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (grover_base / cmim / hybrid / finetuned). Writes one .npy per readout type (atom_from_atom, bond_from_atom, atom_from_bond, bond_from_bond) plus canonical_smiles.npy and validity.npy. Calls task/extract_embeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: targets ~150 lines / ~1800 tokens — within the
# 500-line / 5000-token cap for skill files.
---
# kermt-embed
Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.
The skill is the workflow orchestrator: validate ckpt, validate CSV, clean
SMILES, launch the runner blocking, return the per-readout `.npy` files.
## Hardware requirements
- **GPUs**: 1 (single-GPU).
- **VRAM**: ≥ 4 GB for the default `batch_size 64`.
- **Disk**: depends on output size — roughly a few MB per 1k molecules at
`hidden 800` per readout, so ~10–20 MB per 1k molecules across the 4
readouts. Plus a small `canonical_smiles.npy` + `validity.npy` per run.
- **Driver / CUDA**: any host supporting CUDA 12.6.
## Inputs
Required:
- `--csv <path>` — SMILES CSV. First column is `smiles`; other columns
are ignored (no targets needed).
Checkpoint (optional — defaults to the released model if omitted):
- `--ckpt <path>` — any encoder-bearing checkpoint. Grover_base, cmim,
hybrid, and finetuned ckpts are all accepted. The validator only refuses
ckpts with no encoder. **If omitted**, the skill offers to download the
released pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and embed with
it — see "Resolve & validate the checkpoint" (workflow step 3).
- `--pretrained-release` — explicit opt-in to use the released model without
the interactive prompt (for non-interactive / agent runs). Mutually
exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
`$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
reused, not re-downloaded.
Optional:
- `--batch-size N` — override the configured default (64).
- `--gpus 0` — single GPU id (default 0).
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
`prepare_data.json` in `<dir>`.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout.
1. **Pre-flight: container + system probe.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh check_system
```
2. **Compute run directory.**
```
RUN_DIR=$KERMT_REPO/runs/embed_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Resolve & validate the checkpoint.**
**Resolve — only if `--ckpt` was omitted.** Default to the released
pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
- **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
(NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
and embed with it? [y/N]". **Never download without an explicit yes** (or
`--pretrained-release`). If both `--ckpt` and `--pretrained-release` are
given, abort — they conflict.
- **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
`--model-dir <dir>` if given. An already-complete bundle is reused.
- **Download** (foreground; ~282 MB on first fetch):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
"python agent/scripts/fetch_released_model.py --out /model"
```
Parse the JSON; abort on `ok: false` (surface `errors`). On success set
`<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.
**Validate** the resolved (or user-provided) ckpt:
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
"python agent/scripts/check_checkpoint.py --mode embed --ckpt /ckpt"
```
Parse JSON. Abort on `ok: false`. The validator only refuses encoder-less
ckpts (rare).
4. **Validate the data.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
"python agent/scripts/check_data.py --mode embed --csv /data/<basename>"
```
5. **Prepare the data** (clean-only — no features step).
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode embed \\
--csv /data/<basename> --out /runs/data"
```
Outputs land at `$RUN_DIR/data/prepare_data.json` with a single `clean_csv`
path. `task/extract_embeddings.py` featurizes from SMILES on the fly.
6. **Launch the runner (blocking).**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run \\
--ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
"python agent/scripts/run_extract_embeddings.py \\
--ckpt /ckpt \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs \\
[--gpus 0 --batch-size N]"
```
7. **Report to the user.**
- Embeddings directory: `$RUN_DIR/out/`
- `atom_from_atom.npy`, `bond_from_atom.npy`,
`atom_from_bond.npy`, `bond_from_bond.npy` (the 4 standard readouts;
each shape `(N_rows, hidden_size)`)
- `metadata.pkl` — pickle of a dict containing `canonical_smiles`
(RDKit-canonicalized SMILES per row), `valid` (boolean per-row: did
RDKit parse it), plus other run metadata.
- Manifest: `$RUN_DIR/run.json`
- Log: `$RUN_DIR/logs/embed.log`
## Hard rules
- **Never download the released model without consent.** When `--ckpt` is
omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
or an explicit `--pretrained-release` flag. `--ckpt` and
`--pretrained-release` are mutually exclusive.
- **Never modify the user's ckpt.** The runner reads-only via
`task/extract_embeddings.py`'s `--checkpoint <path>` flag.
- **Arch comes from the ckpt.** No `--hidden-size` flag etc. on this runner;
`task/extract_embeddings.py` reads arch from the ckpt's saved_args.
## Common errors
- `prepare_data manifest is missing required output 'clean_csv'` → prepare
ran with `--skip-clean` but no source CSV given. Re-run prepare without it.
- `--gpus '0,1' is single-GPU only` → pass a single id.
## Replayability
```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```
If `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin
the commit via `repo.commit` and `git checkout` it first.
kermt-finetune15.2 KB
---
name: kermt-finetune
description: Finetune a pretrained KERMT encoder on a labeled CSV. The skill validates the input checkpoint (must be a pretrain ckpt — grover_base / cmim / hybrid), validates the labeled CSV, prepares the data (clean + features + optional split), then launches main.py finetune inside the kermt container (detached for hours-scale runs). Hyperparameters come from agent/config/defaults_finetune.json with per-flag CLI override.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: targets ~250 lines / ~3000 tokens — within the
# 500-line / 5000-token cap for skill files.
---
# kermt-finetune
Finetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill
is the workflow orchestrator: validate ckpt, validate data, prepare data,
launch the runner detached, return a run directory + container name.
## Hardware requirements
- **GPUs**: 1 by default (single-GPU); pass `--gpus 0` (or whichever id) to
select one. For faster training on a multi-GPU host, pass `--num-gpus N`
(N>1) to run data-parallel DDP across N GPUs — `--batch-size` is then
per-GPU (effective global batch = batch_size × N).
- **VRAM**: ≥ 8 GB for the default `batch_size 32` configuration. Lower VRAM
works at smaller batch sizes — pass `--batch-size N` to override.
- **Disk**: a few GB per run (checkpoint + features + logs).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).
`kermt-setup` validates this up-front.
## Inputs
Required:
- `--csv <path>` — labeled CSV. First column is `smiles`; every other column
is a target.
Checkpoint (optional — defaults to the released model if omitted):
- `--ckpt <path>` — input pretrain checkpoint (grover_base / cmim / hybrid).
The validator refuses already-finetuned ckpts with a redirect to
`kermt-infer`. **If omitted**, the skill offers to download the released
pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and finetune from it —
see "Resolve & validate the checkpoint" (workflow step 3).
- `--pretrained-release` — explicit opt-in to use the released model without
the interactive prompt (for non-interactive / agent runs). Mutually
exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
`$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
reused, not re-downloaded.
Optional:
- `--dataset-type {regression | classification | multiclass}` — default
`regression` (from `defaults_finetune.json`). Drives loss, metric defaults,
and head initialization. For classification tasks pass
`--dataset-type classification`.
- `--targets COL [COL ...]` — explicit target column names. If omitted, the
validator auto-detects numeric non-smiles columns and the skill confirms
with the user before proceeding.
- `--val-csv <path>` and `--test-csv <path>` — user-provided val + test
splits. Either pass both or pass neither (the skill auto-splits using the
configured `--split-type`).
- `--split-type {random | scaffold_balanced | index_predetermined}` —
default `scaffold_balanced` from `defaults_finetune.json`.
- `random` and `scaffold_balanced`: build the val/test split internally
from the train CSV. No `--val-csv` / `--test-csv` needed.
- `index_predetermined`: **requires** pre-split CSVs passed via
`--val-csv` + `--test-csv` (and, separately, per-fold index files —
see `kermt/util/utils.split_data`). Use this when the dataset ships
its own canonical split (e.g. `tests/data/Biogen_for_grover/scaffold/
balance/<endpoint>/{train,val,test}.csv`).
- `--metric NAME` — `mae` (regression default), `auc` (classification default),
or any name `kermt.util.metrics.get_metric_func` accepts.
- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /
`--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /
`--bond-drop-rate F` / `--dist-coff F` / `--early-stop-epoch N` /
`--seed N` — training-hyperparameter overrides. Anything not given is
filled from `agent/config/defaults_finetune.json`.
- `--ffn-hidden-size N` / `--ffn-num-layers N` — shared FFN trunk dims.
- `--ffn-num-task-specific-layers N` / `--ffn-task-specific-hidden-size H` —
per-target FFN heads (default 0 = off; useful for heterogeneous multi-target
finetunes). Both must be set together when N > 0.
- `--ensemble-size N` / `--num-folds N` — multi-model / k-fold CV. Default 1
each.
- `--gpus 0` — single GPU id for single-process finetune (default 0). Ignored
when `--num-gpus > 1`.
- `--num-gpus N` — number of GPUs for data-parallel DDP finetune. Default 1
(single-process, unchanged). N>1 runs `main.py finetune` with `WORLD_SIZE=N`
(one process per GPU); `--batch-size` is per-GPU.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
`prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has built `kermt:latest`. All paths below are on the host; the
helper bind-mounts them at known container paths.
1. **Pre-flight: ensure container + system probe.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c "
import json, sys; d = json.load(sys.stdin)
if not d['ok']:
print('System check failed:', d['gaps']); sys.exit(1)
print(f'OK: {len(d[\"gpus\"])} GPU(s); CUDA via container toolkit')
"
```
Refuse to proceed if `ok: false`.
2. **Compute run directory.**
```
RUN_DIR=$KERMT_REPO/runs/finetune_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Resolve & validate the checkpoint.**
**Resolve — only if `--ckpt` was omitted.** Default to the released
pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
- **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
"No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
(NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
and finetune from it? [y/N]". **Never download without an explicit yes**
(or `--pretrained-release`). If both `--ckpt` and `--pretrained-release`
are given, abort — they conflict.
- **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
`--model-dir <dir>` if given. An already-complete bundle is reused.
- **Download** (foreground; ~282 MB on first fetch):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
"python agent/scripts/fetch_released_model.py --out /model"
```
Parse the JSON; abort on `ok: false` (surface `errors`). On success set
`<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.
**Validate** the resolved (or user-provided) ckpt:
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
"python agent/scripts/check_checkpoint.py --mode finetune_init --ckpt /ckpt"
```
Parse the JSON. Abort on `ok: false`. The validator rejects already-
finetuned ckpts (`has_task_ffn: true`) with a redirect to `kermt-infer`.
4. **Validate the data.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
"python agent/scripts/check_data.py --mode finetune --csv /data/<basename> [--targets COL1 COL2 ...]"
```
If `--targets` was not given by the user, surface `auto_detected_targets`
from the JSON and ask the user to confirm before continuing. Abort on
`ok: false`.
5. **Prepare the data** (skip if `--from-prepare` given).
**Pre-flight: check for sibling val.csv / test.csv.** Before invoking
prepare_data, inspect the parent directory of `<user-csv>`. If a
canonical-looking sibling `val.csv` (or `val_*.csv` — common variants
include `val_T.csv`, `val_clean.csv`) AND a matching `test.csv` /
`test_*.csv` exist next to the train CSV, the dataset ships its own
pre-defined split. **In that case set `--split-type index_predetermined`
AND pass `--val-csv` / `--test-csv`** — otherwise the configured
`split_type` (default `scaffold_balanced`) will re-split the train CSV
from scratch and silently discard the user's val/test files. When in
doubt — or when the sibling files use non-canonical suffixes (`_T`,
`_v2`, etc.) — surface the situation to the user and ask which they
want.
**Quoting target names.** If any of the `--targets` column names
contain shell metacharacters (`>`, `&`, `|`, `(`, `)`, `$`, etc.),
single-quote each one when passing on the CLI to keep the shell from
eating part of the name. Example: `--targets 'Log_Caco2_Papp_A>B'
'logD'`. The CSV header itself is read directly by the downstream
trainer and is unaffected, but the prepare_data.json manifest's
`targets[]` field captures whatever the shell delivers — unquoted
metacharacters get truncated there.
**Mount note:** `kermt_container.sh --data <host-csv>` mounts the
parent directory of `<host-csv>` at `/data`. `--val-csv` and
`--test-csv` must therefore reference files in that same parent
directory. If val/test live in a separate directory (e.g. a sibling
`splits/` folder), mount the parent of all three using `--data <dir>`
on a directory rather than a file.
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode finetune \\
--csv /data/<basename> --out /runs/data \\
--split-type <split_type> \\
[--val-csv /data/<val-basename> --test-csv /data/<test-basename>] \\
[--val-frac 0.1 --test-frac 0.1 --seed 0] \\
--targets <COL1> [COL2 ...]"
```
Outputs land at `$RUN_DIR/data/prepare_data.json`. For `scaffold_balanced`
and `index_predetermined`, prep emits a single `clean_full_csv` + `.npz`;
the runner passes them through to `main.py finetune` which calls
`split_data` internally with the user-supplied seed.
6. **Estimate runtime + echo applied defaults.**
- Finetune wall time is typically minutes-to-hours on 1 GPU.
- Surface a summary of every flag that was filled from the defaults
vs user-supplied, so the user knows what was assumed. The runner
records this in `args_applied`.
- Sample message:
`"Filling from defaults_finetune.json: epochs=30, batch_size=32,
split_type=scaffold_balanced. Override any of these with --<flag>."`
7. **Targets confirmation gate (hard requirement).** Before launching the
runner, regardless of how the targets list was determined (CLI `--targets`,
auto-detection in step 4, or a user natural-language request like
"finetune on Caco2 and HLM"), echo the final targets list to the user with
an explicit count:
`"Will finetune on N target(s): COL1, COL2, ..."`. If the user's request
specified a subset that doesn't match this list (e.g., they asked for 2
tasks via natural language but the list still has 4), treat it as a
discrepancy and re-prompt with the diff — never silently proceed on the
wrong target set. Wait for explicit confirmation before launching unless
`--yes` was given.
8. **Launch the runner detached.** (Consistent with the pretrain skills.)
```
$KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
--name kermt-finetune-<ts> \\
--ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
"python agent/scripts/run_finetune_local.py \\
--ckpt /ckpt \\
--prepare-manifest /runs/data/prepare_data.json \\
--dataset-type <type> \\
--out /runs \\
[--gpus 0] \\
[--num-gpus N] \\
[--epochs N --batch-size N --init-lr F ...] \\
[--ffn-num-task-specific-layers N --ffn-task-specific-hidden-size H]"
```
Returns the container name + id + log file path.
9. **Report to the user.** Output a short summary:
- Container name + id
- `$RUN_DIR/run.json` (manifest with cmd_replay + image digest)
- Log file: `$RUN_DIR/logs/finetune.log`
- TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
$RUN_DIR/logs/tb`)
- Final checkpoints land at `$RUN_DIR/ckpt/fold_0/model_0/model.pt`
(best-val) and `last_checkpoint.pt` (sibling, auto-resume target).
Held-out test predictions + metrics land at
`$RUN_DIR/ckpt/fold_0/test_result.csv`. Paths vary with `--num-folds`
/ `--ensemble-size`.
- To follow progress: `kermt-monitor <RUN_DIR>` (one-shot) or
`docker logs -f <container-name>` (streaming).
- To block until the run finishes (useful for short test runs):
`docker wait <container-name>` — prints the exit code on completion.
## Hard rules
- **Never download the released model without consent.** When `--ckpt` is
omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
or an explicit `--pretrained-release` flag. `--ckpt` and
`--pretrained-release` are mutually exclusive.
- **Never modify the user's input ckpt.** The runner passes its path via
`--checkpoint_path`; `task/train.py` loads it read-only into the model and
attaches a new FFN head. The source file stays untouched.
- **Arch comes from the ckpt, not from CLI/defaults.** The runner extracts
`hidden_size`, `depth`, `num_attn_head`, `activation`, `embedding_output_type`,
`self_attention` (+ `attn_hidden` / `attn_out` when applicable) from the
ckpt's saved_args. There is no `--hidden-size` flag on this runner.
- **Never block on the long-running finetune.** The skill launches via
`run_detached` and returns immediately after step 9. Use `kermt-monitor`.
- **Echo applied defaults back to the user.** The `args_applied` field of
`run.json` records every flag's value + source (user / default-config).
Surface a one-line summary of every filled-from-default flag so the user
knows what was assumed.
## Common errors
- `finetune_init requires a pretrain ckpt (grover_base / cmim / hybrid)` →
the ckpt you passed is already finetuned (has task FFN heads). Pick a
pretrain ckpt instead, or use `kermt-infer` if you want to run
predictions with the existing finetuned model. To resume a finetune on
the SAME dataset, bypass the skill and call
`python main.py finetune --checkpoint_path <ckpt> ...` directly — the
agent skill doesn't support resume because saved-task identity
can't be machine-verified against the new training data.
- `prepare_data manifest reports ok=False` → check `errors` for the failed
step (typically clean_smiles or save_features). Fix and re-run.
- `ffn_num_task_specific_layers=N>0 but ffn_task_specific_hidden_size is unset`
→ MTL heads need an explicit hidden size. Pass `--ffn-task-specific-hidden-size H`.
- `finetune is single-GPU` (from `--gpus 0,1`) → `--gpus` selects one device
for single-process finetune. For multi-GPU, use `--num-gpus N` (DDP) instead.
## Replayability
The `run.json` `cmd_replay` field is a single-line command that re-runs the
finetune with the same inputs, hyperparameters, and arch. To replay inside
the kermt container:
```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```
If `ok_to_replay: false` in the manifest (because the kermt repo working
tree was dirty at launch time), the replay may not be bit-exact — pin the
exact commit via the `repo.commit` field and `git checkout` it
first.
kermt-infer5.56 KB
---
name: kermt-infer
description: Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The skill validates that the input ckpt has task FFN heads (refuses pretrain ckpts with a redirect to kermt-finetune), validates the CSV, prepares the data (clean + rdkit_2d features), then launches main.py predict inside the kermt container (blocking, minutes-scale).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: targets ~170 lines / ~2000 tokens — well within the
# 500-line / 5000-token cap for skill files.
---
# kermt-infer
Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The
skill is the workflow orchestrator: validate ckpt, validate CSV, prepare data,
launch the runner blocking, return the predictions CSV.
## Hardware requirements
- **GPUs**: 1 (single-GPU). Multi-GPU inference is not currently supported.
- **VRAM**: ≥ 4 GB for the default `batch_size 32`.
- **Disk**: a few hundred MB per run (cleaned CSV + features + predictions).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).
## Inputs
Required:
- `--ckpt <path>` — finetuned checkpoint (must have task FFN heads). The
validator refuses pretrain ckpts with a redirect to `kermt-finetune`.
- `--csv <path>` — SMILES-only CSV. First column is `smiles`; other columns
are ignored.
Optional:
- `--batch-size N` — override the configured default (32).
- `--seed N` — random seed for inference (deterministic featurization paths).
- `--gpus 0` — single GPU id (default 0). Multi-GPU rejected.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
`prepare_data.json` in `<dir>`.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has built `kermt:latest`.
1. **Pre-flight: ensure container + system probe.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh check_system
```
Refuse to proceed on `ok: false`.
2. **Compute run directory.**
```
RUN_DIR=$KERMT_REPO/runs/infer_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Validate the checkpoint.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
"python agent/scripts/check_checkpoint.py --mode inference --ckpt /ckpt"
```
Parse the JSON. Abort on `ok: false`. The validator rejects pretrain ckpts
(`has_task_ffn: false`) with a redirect to `kermt-finetune`.
4. **Validate the data.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
"python agent/scripts/check_data.py --mode inference --csv /data/<basename>"
```
Abort on `ok: false`.
5. **Prepare the data.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode inference \\
--csv /data/<basename> --out /runs/data"
```
Outputs land at `$RUN_DIR/data/prepare_data.json` with `clean_csv` +
`clean_npz` paths (rdkit_2d_normalized features).
6. **Launch the runner (blocking).**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run \\
--ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
"python agent/scripts/run_inference.py \\
--ckpt /ckpt \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs \\
[--gpus 0 --batch-size N --seed N]"
```
Returns the predictions CSV path on success.
7. **Report to the user.** Output a short summary:
- Predictions: `$RUN_DIR/out/predictions.csv` (smiles + per-target columns)
- Manifest: `$RUN_DIR/run.json` (cmd_replay + image digest + applied args)
- Log: `$RUN_DIR/logs/inference.log`
- Row count: <N> molecules predicted across <K> targets
## Hard rules
- **Never modify the user's ckpt.** The runner symlinks the ckpt into a
unique `<out>/ckpt_link/` subdir so `main.py predict --checkpoint_dir`
picks it up; the source file stays untouched.
- **Arch comes from the ckpt, never from CLI/defaults.** The runner records
the validator's arch block in `run.json` but does not pass arch flags into
`main.py predict` — predict reads them from the loaded ckpt's saved_args.
- **Single-GPU only.** Multi-GPU inference is not currently supported.
- **Echo applied defaults.** The `args_applied` field of `run.json` records
every flag's value + source (user / default-config). Surface a short
summary of any default-filled flag.
## Common errors
- `inference requires a finetuned ckpt with task FFN heads` → ckpt is a
pretrain ckpt; use `kermt-finetune` first.
- `prepare_data manifest reports ok=False` → check the manifest `errors` for
the failed step (typically clean_smiles or save_features).
- `could not convert string to float: '<value>'` from save_features or main.py
predict → input CSV has a non-numeric passthrough column (e.g. a 'split'
label). The prep step now strips the CSV to SMILES-only at inference; if
this error still surfaces, the CSV is being read by a runner that bypassed
prepare_data. Re-run via the skill, not `main.py` directly.
- `--gpus '0,1' is single-GPU only` → pass a single id.
## Replayability
The `run.json` `cmd_replay` field is a single-line command that re-runs the
inference with the same inputs. To replay inside the kermt container:
```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```
If `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin
the commit via `repo.commit` and `git checkout` it first.
kermt-monitor6.95 KB
---
name: kermt-monitor
description: Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).
license: Apache-2.0
compatibility: Requires docker and jq. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: atomic-skill
risk_tier: skill
# Line/token budget: this file is targeted at ~120 lines / ~1500 tokens — well
# within the 500-line / 5000-token cap for skill files.
---
# kermt-monitor
Companion skill for any KERMT workflow that runs detached: the three pretrain
skills (`kermt-continue-pretrain`, `kermt-pretrain-scratch`,
`kermt-add-cmim-pretrain`) plus `kermt-finetune`. `kermt-infer` and
`kermt-embed` run blocking by default and don't need this skill, but if a
user launches them detached on purpose the monitor still works (the
workflow-dispatch in step 4 handles unknown workflows by tailing the
most-recent log file in the run dir). Reads the run directory's `run.json`,
queries docker for the container's state, surfaces the latest progress,
and either tails or follows the log.
## Hardware requirements
None. This skill only reads disk + queries docker; no GPU compute.
## Inputs
One of:
- `<run-dir>` — a positional argument pointing at the directory containing
`run.json` (e.g. `runs/continue-pretrain_2026-05-17T10-23Z`). Preferred.
- `--container <name-or-id>` — direct container reference; the skill still
reads `run.json` from the run dir referenced inside the container's
inspect output if available, but works degraded-mode without it.
Optional:
- `--lines N` — number of trailing log lines to print (default 50).
- `--follow` — stream `docker logs -f` until ^C. Useful for "watch the
loss". Without it, the skill is one-shot and exits.
- `--json` — emit a structured status report instead of human-readable text.
Useful when the parent agent wants to take downstream action.
## Workflow
Let `RUN_DIR=$1` (or whatever path the user supplies).
1. **Locate the manifest.**
```
MANIFEST=$RUN_DIR/run.json
```
Refuse to proceed if it doesn't exist; surface a helpful message
pointing the user at the run-dir convention (`runs/<workflow>_<ts>/`).
2. **Parse the manifest** (Python helper):
```
workflow=$(jq -r .workflow $MANIFEST)
container_name=... # not directly in run.json today; the skill that
# launched stored it in run.json under
# container.name during launch (see below note).
logs_dir=$(jq -r .logs_dir $MANIFEST)
image_tag=$(jq -r .container.image_tag $MANIFEST)
started_at=$(jq -r .started_at $MANIFEST)
```
3. **Query docker for container state.**
```
docker ps --filter "name=$container_name" --format \
'{{.ID}}\t{{.Status}}\t{{.CreatedAt}}'
```
If absent, fall back to `docker inspect $container_name --format
'{{.State.Status}} (exit {{.State.ExitCode}})'` to see whether the
container exited (ok or failed) or was removed (`--rm` after exit).
4. **Find the live log file.**
```
case "$workflow" in
continue-pretrain|pretrain-scratch) LOG=$logs_dir/pretrain_ddp.log ;;
finetune) LOG=$logs_dir/finetune.log ;;
*) LOG=$(ls -1t $logs_dir/*.log 2>/dev/null | head -n 1) ;;
esac
```
The manifest's `workflow` field disambiguates pretrain (`pretrain_ddp.log`)
from finetune (`finetune.log`). Other workflows fall back to the
most-recently-modified `.log` in `$logs_dir`.
5. **Show the latest progress.**
- `tail -n $LINES $LOG` for the raw recent output.
- Parse the last few progress lines and surface a human-friendly
summary. The format differs per workflow:
- Pretrain: epoch / step / val_loss
```
Current epoch: 12/100 step: 4523/9000 val_loss: 0.832 (best 0.821 @ step 4100)
```
- Finetune: fold / epoch / val_<metric> (e.g. val_mae for regression,
val_auc for classification — read `args_applied.metric` from run.json)
```
Fold 0 epoch 12/30 val_mae 0.187 (best 0.182 @ epoch 9)
```
```
Wall-clock: 1h 23m since started_at; ETA ~6h remaining.
```
6. **Final test-metrics block (finetune, on completion).** If `workflow` is
`finetune` AND the container has exited cleanly (`State.Status=exited`,
`ExitCode=0`) AND `$RUN_DIR/ckpt/fold_*/test_result.csv` exists, parse it
and emit a per-task metric table:
```
Final test metrics (per task):
Target MAE
HLM_clearance 0.187
RLM_clearance 0.213
MDR1-MDCK_efflux 0.241
solubility_pH6.8 0.156
```
The metric column matches `args_applied.metric` (mae for regression, auc
for classification, etc.). For multi-fold or ensemble runs, average across
folds/models and note `± std` if std > 0. Skip silently if no
`test_result.csv` exists (run incomplete or no test split was emitted).
7. **If `--follow`, stream live logs.**
```
docker logs -f $container_name
```
Wraps until ^C.
8. **Stop / cleanup hints** (printed at end of one-shot mode):
```
To stop: docker stop $container_name
To remove: docker rm $container_name
To re-run: `$(jq -r .cmd_replay $MANIFEST)`
```
## Hard rules
- **Read-only on the user's data.** Never modify `run.json`, never touch the
container's checkpoint dir. The monitor only inspects.
- **Don't kill the container without explicit user instruction.** If the
user asks to stop, run `docker stop`; if they ask to abandon, leave it
running and just exit.
- **Don't pull or modify the kermt image.** The monitor only reads.
- **JSON output mode is non-interactive.** Skip the "press ^C to exit"
prompts and emit a single JSON document so the parent agent can pipe it.
## Note on container_name plumbing
The run.json schema as currently written does not yet include the launched
container name — `kermt_run_detached` prints it to stdout but the runner
script doesn't capture it into run.json. The monitor falls back to a
filesystem-based lookup: list `runs/<workflow>_*/` directories and match by
mtime; or accept `--container <name>` explicitly. Follow-up: have the
launching skill record container name into run.json before exiting.
## Output (text mode, default)
```
KERMT continue-pretrain · runs/continue-pretrain_2026-05-17T10-23Z
Container : kermt-continue-pretrain-… (Up 1 hour, status: running)
Image : kermt:latest@sha256:…
Repo : 2fe00f9 (clean)
Started : 2026-05-17T10:23:14Z (1h 23m ago)
Workflow : continue-pretrain, pretrain_mode=hybrid, world_size=2
Latest log (last 50 lines from $LOG):
[Epoch 12/100] step 4523/9000 loss 0.832 lr 1.2e-4
[val] step 4100 val_loss 0.821 (new best)
...
Progress: epoch 12/100, ~12% done. ETA ~6h.
TensorBoard: tensorboard --logdir $RUN_DIR/logs/tb
Replay command: $(jq -r .cmd_replay $RUN_DIR/run.json)
```
kermt-pretrain-scratch8.99 KB
---
name: kermt-pretrain-scratch
description: Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: workflow-skill
risk_tier: skill
# Line/token budget: ~210 lines, ~2400 tokens — well within the
# 500-line / 5000-token cap for skill files. Most of the orchestration is
# shared with kermt-continue-pretrain; the differences are documented below.
---
# kermt-pretrain-scratch
Pretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful
when you want to retrain a model on a custom chemistry domain rather than
extending one of the released checkpoints. **Significantly more expensive than
`kermt-continue-pretrain`** — no warm start, so the loss curves need to descend
from scratch over many epochs.
## Hardware requirements
Same as `kermt-continue-pretrain`:
- **GPUs**: 1–N CUDA-capable. The runner auto-detects via
`torch.cuda.device_count()`; `--gpus 0,2` overrides. Single-GPU fallback:
`--batch_size 32 --save_interval 500`. Multi-GPU keeps defaults
(`--batch_size 256` etc.). Note: `--gpus N` uses **torch.cuda** indexing,
which can differ from `nvidia-smi`'s display order on multi-GPU hosts
(PCI bus vs. CUDA enumeration). To target a specific physical GPU, set
`CUDA_VISIBLE_DEVICES` before invoking, or run
`python -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])"`
to confirm which device you're picking.
- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware
(80 GB VRAM). On smaller GPUs, downscale to avoid OOM:
| GPU class | VRAM | Suggested `--batch-size` |
|--------------------------|------------|--------------------------|
| L4, T4, V100 16 GB | 16–24 GB | 32–64 |
| A100 40 GB, L40, A40 | 40–48 GB | 128 |
| A100 80 GB, H100, H200 | 80 GB | 256 (default) |
These are rough starting points — pass `--batch-size N` to override.
- **Disk**: tens of GB for shards + vocab + checkpoints, scaled by epochs.
- **Wall time**: this is the big difference. Pretraining from scratch on an
11M-mol corpus at 100 epochs typically takes **days even on a multi-GPU box**.
The skill prints an estimate before launching; confirm with the user.
## When to invoke
- User wants to train a new model on a custom corpus (e.g. domain-specific
chemistry that the released ckpts don't cover).
- User wants to reproduce a pretrain config end-to-end without depending on a
released ckpt.
For continuing an existing released ckpt, use `kermt-continue-pretrain`. For
adding a cMIM decoder to an encoder-only grover_base ckpt, use
`kermt-add-cmim-pretrain`.
## Inputs
Required:
- `--csv <path>` — the pretrain corpus CSV with a `smiles` column. Single file
by convention; multi-file corpora deferred. Use `--val-csv` for a separate
validation set.
- `--pretrain-target-mode {vocab|cmim|hybrid}` — which pretrain objective to
use. **No default** — must be set explicitly so the user makes an informed
choice:
- `vocab` — original GROVER-style atom + bond vocab prediction (encoder-only
output, lightweight).
- `cmim` — contrastive + SMILES reconstruction objective. Requires building
a SMILES vocab from the corpus.
- `hybrid` — both vocab and contrastive objectives jointly (the
state-of-the-art config from the KERMT manuscript).
Optional:
- `--val-csv <path>` — separate validation CSV. Without it, prepare_data
auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).
- Training-hyperparameter overrides: `--epochs N` / `--batch-size N` /
`--init-lr F` / `--max-lr F` / `--final-lr F` / `--warmup-epochs F` /
`--weight-decay F` / `--dropout F` / `--save-interval N` / `--seed N`.
Anything not given is filled from `agent/config/defaults_pretrain.json`.
- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /
`--contrastive-temperature F` (cmim and hybrid only).
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run
name is honored only alongside a project. Off by default.
- `--gpus 0,2` — restrict to a GPU subset.
## Workflow
Let `$KERMT_REPO` be the path to your kermt repo checkout.
1. **Pre-flight: ensure container + system probe** (same as
`kermt-continue-pretrain` step 1). Refuse to proceed if `check_system`
reports gaps.
2. **Compute run directory.**
```
RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ)
```
3. **Validate the corpus** (no ckpt to validate, so this is the only input
check):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
"python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>"
```
Abort on `ok: false`.
4. **Prepare the data** — no vocab pass-through (we want fresh vocab from
corpus):
```
$KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
"python agent/scripts/prepare_data.py --mode pretrain \\
--csv /data/<basename> --out /runs/data \\
[--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
```
Outputs land at `$RUN_DIR/data/prepare_data.json` with
`vocab_source: "built_fresh"`.
5. **Estimate runtime + warn loudly.** This is critical for pretrain-from-scratch:
- "Pretraining from scratch is days-scale even on multi-GPU; the released
KERMT checkpoints were each trained on millions of molecules for hundreds
of GPU-hours. If you mainly want to leverage existing knowledge for a
downstream task, consider `kermt-continue-pretrain` from a released ckpt
instead, which converges in hours instead of days."
- Show the corpus size × epochs × GPU count → estimated wall time.
- Ask for explicit confirmation unless `--yes` was given.
6. **Launch the runner detached.**
```
$KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
--name kermt-pretrain-scratch-<ts> \\
--run-dir $RUN_DIR -- \\
"python agent/scripts/run_pretrain_local.py \\
--from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\
--prepare-manifest /runs/data/prepare_data.json \\
--out /runs \\
[--epochs N --batch-size N ...]"
```
Note: NO `--ckpt` flag (the runner refuses if both `--from-scratch` and
`--ckpt` are given). The runner uses the `arch` group from
`agent/config/defaults_pretrain.json` to size the model.
7. **Report to the user.** Always include all of the following — do not
omit the TensorBoard line under output-length pressure:
- Container name + id
- `$RUN_DIR/run.json` (the manifest with `workflow: pretrain-scratch`,
`from_scratch: true`, `vocab_check: null`, `arch` from defaults, full
`cmd_replay`)
- Log file: `$RUN_DIR/logs/pretrain_ddp.log`
- TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
$RUN_DIR/logs/tb`)
- Suggest `kermt-monitor <RUN_DIR>` for progress.
## Hard rules
- **Never accept a `--ckpt` flag.** From-scratch is exclusive with input
ckpt — the runner enforces this; the skill should too.
- **Never silently default `--pretrain-target-mode`.** This is a significant
architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user
if not given on the CLI.
- **Strong warning before launching.** From-scratch pretrain is the most
expensive workflow. The user needs to know what they're committing to.
## Common errors
- `--pretrain-target-mode is required when --from-scratch is set` → user
forgot the mode flag. Prompt.
- `--from-scratch is incompatible with --ckpt` → user provided both; ask which
one they meant.
- `defaults_pretrain.json has no arch group` → repo state issue (should never
happen on a fresh clone); points the user at running `kermt-setup` again.
## What's in the manifest after a from-scratch run
Same reproducibility fields as continue-pretrain (`repo.commit`, `kermt_image`,
`cmd_replay`, `args_applied`), plus:
- `workflow`: `"pretrain-scratch"`
- `from_scratch`: `true`
- `inputs.ckpt`: `null`
- `ckpt_symlink`: `null`
- `vocab_check`: `null` (not verified — vocab built from corpus is
authoritative for from-scratch)
- `arch`: the values pulled from `agent/config/defaults_pretrain.json`'s
`arch` group (with any future CLI overrides applied).
## Replayability
Same as continue-pretrain: `cmd_replay` is a copy-pasteable command. If
`ok_to_replay: false`, the kermt repo working tree was dirty at launch
time — check `repo.commit` and `git checkout` it first.
kermt-setup6.18 KB
---
name: kermt-setup
description: Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test inside the container. Every other kermt-* skill depends on this; invoke it first.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
owner: evax@nvidia.com
classification: atomic-skill
risk_tier: skill
# This file is intentionally short (~110 lines, ~1200 tokens) — well within the
# 500-line / 5000-token budget for skill files. Longer reference material lives
# alongside agent/scripts/kermt_container.sh.
---
# kermt-setup
Bootstrap the KERMT agent environment. Run this once on a fresh machine (or
after the Dockerfile or `environment.yml` changes) before invoking any other
`kermt-*` skill.
## Hardware requirements
- **GPU**: at least one CUDA-capable NVIDIA GPU visible to the host. The image
is based on `nvidia/cuda:12.6.3-cudnn-devel-ubuntu22.04`, so the host driver
must support CUDA 12.6. Verify with host `nvidia-smi` before invoking.
- **Host docker**: docker engine + nvidia-container-toolkit. Without the
toolkit, `docker run --gpus all` will fail at step 2 of the workflow below.
- **Disk**: ≈ 50 GB free for the built kermt image (`docker image inspect
--format '{{.Size}}'` reports ≈ 44 GB; the `docker images` Size column
can show ~100 GB because it counts shareable buildx attestation layers
that are deduplicated across images). Plan for ~50 GB of unique on-disk
storage; add a comfortable buffer if you're also keeping build cache.
- **Memory**: the build itself peaks at ~4 GB RAM during conda env solve.
- This skill does not run training/inference workloads itself; per-workflow
hardware requirements (VRAM, GPU count) are declared in the respective
`kermt-<workflow>` skills.
## When to invoke
- User explicitly asks (`/kermt-setup`, "set up kermt", "build the kermt image",
etc.).
- Or another `kermt-*` skill detected that the image does not exist and routed
here. (Most other skills call `kermt_ensure_image` themselves, so this is
usually only needed for the first-time setup, debugging, or a forced rebuild.)
## Inputs
The skill takes no required arguments. Optional overrides (via env vars before
invoking, or by setting them in the user's shell):
- `KERMT_IMAGE` — image tag to build/verify (default: `kermt:latest`).
- `KERMT_REPO` — host path of the kermt repo checkout (default: auto-derived
from the script's location).
If the user has not specified a repo path and the current working directory is
not inside a kermt repo clone, ask for the repo path before proceeding.
## Workflow
All work goes through `agent/scripts/kermt_container.sh`. The script's
subcommand dispatch can be invoked directly without sourcing — that is the
preferred form for skill use.
Let `HELPER=$KERMT_REPO/agent/scripts/kermt_container.sh`.
1. **Verify docker is installed and the daemon is reachable.**
```
$HELPER check_docker
```
Exit 0 → continue. Non-zero → surface the error to the user (typically
"docker not on PATH" or "daemon not reachable"); do not attempt step 2.
2. **Verify GPU passthrough works.**
```
$HELPER check_gpu
```
This runs `docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04
nvidia-smi` and checks the exit status. Non-zero → tell the user to install
`nvidia-container-toolkit` on the host and confirm a CUDA-capable NVIDIA GPU
is visible to the host (`nvidia-smi` on the host should also work). Stop
here; without GPU passthrough the kermt image will build but no workflow
will run.
3. **Build or verify the kermt image.**
```
$HELPER ensure_image
```
If the image already exists, this returns immediately. Otherwise it builds
from `$KERMT_REPO/Dockerfile`. **Warn the user before invoking** that the
first build takes ~10–20 minutes on a typical workstation and streams build
logs to the console. Do not run this in the background — the user wants to
see progress and any build failures must surface immediately.
4. **GPU smoke test inside the container.** Quote the whole `python` command
as a single string — the helper passes args through `bash -c "$*"`, so
unquoted multi-word commands get re-parsed and any embedded quotes are
collapsed.
```
$HELPER run -- 'python -c "import torch; print(\"cuda_available:\", torch.cuda.is_available()); print(\"device_count:\", torch.cuda.device_count())"'
```
Expected output: `cuda_available: True` and a positive `device_count`. If
`cuda_available` is `False` despite step 2 passing, something is wrong with
the container's CUDA wiring — report the full output to the user and stop;
do not declare the environment ready.
5. **Summary to user.** Report:
- Image tag and ID (`docker image inspect $KERMT_IMAGE --format '{{.Id}}'`).
- Image size (`docker image inspect $KERMT_IMAGE --format '{{.Size}}'`).
- GPU count detected inside the container.
- "Ready" — the user can now invoke other `kermt-*` skills.
## Hard rules
- Do **not** pull or push docker images. The kermt image is built locally only.
- Do **not** auto-delete or prune older `kermt:*` tags without the user's
explicit confirmation — the user may be running a finetune or pretrain in
another container that depends on a specific tag.
- Do **not** modify the host's docker daemon configuration, daemon.json, or
user-group membership.
- Do **not** modify the `Dockerfile` or `environment.yml` as part of this
skill. If the build fails because of a Dockerfile issue, surface the error
and stop; let the user decide whether to edit.
- Do **not** rebuild the image when it already exists (i.e. do not pass a
`--no-cache` or `--pull` flag to ensure_image) unless the user explicitly
asks for a forced rebuild.
## Forced rebuild
If the user explicitly asks to rebuild (e.g. after changing the Dockerfile or
`environment.yml`), the cleanest path is to remove the old image first, then
rerun `ensure_image`:
```
docker image rm $KERMT_IMAGE
$HELPER ensure_image
```
Confirm with the user before running `docker image rm`.
molmim-nim8.23 KB
---
name: molmim-nim
description: >
Use this skill for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization. Invoke for MolMIM, molecular embeddings, hidden states, latent decoding, sampling around a seed SMILES, CMA-ES guided molecule generation, QED or plogP optimization, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28; rdkit"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# MolMIM NIM
Generate, sample, embed, and decode small molecules with MolMIM. Use this
`SKILL.md` for first-pass hosted/local usage; load supplemental files only when
needed:
- `references/api.md`: endpoints, schema, Docker flags, response fields.
- `references/science.md`: use cases, strengths, limits, and handoffs.
- `references/parameters.md`: generation, sampling, and optimization effects.
- `references/validation.md`: SMILES/property/artifact checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted generation: `https://health.api.nvidia.com/v1/biology/nvidia/molmim/generate`
- Local generation: `http://localhost:8000/generate`
- Local embedding: `http://localhost:8000/embedding`
- Local hidden state: `http://localhost:8000/hidden`
- Local decode: `http://localhost:8000/decode`
- Local sampling: `http://localhost:8000/sampling`
- Local readiness: `http://localhost:8000/v1/health/ready`
Mode difference: the hosted API reference exposes `/generate`; the local
container exposes the broader latent-space workflow (`/embedding`, `/hidden`,
`/decode`, `/sampling`, `/generate`). Do not invent hosted latent endpoints.
Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Local inference uses
no auth header after readiness.
## Local Docker
Use shell env first; source repo-root `.env` only if present. Do not print keys.
MolMIM docs use `NGC_CLI_API_KEY` for the local container; this repo accepts
`NGC_API_KEY` or `NVIDIA_API_KEY` and maps to `NGC_CLI_API_KEY` for startup.
Mount `LOCAL_NIM_CACHE` at `/home/nvs/.cache/nim`.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
if [ -z "${NGC_CLI_API_KEY:-}" ] && [ -n "${NGC_API_KEY:-}" ]; then
export NGC_CLI_API_KEY="$NGC_API_KEY"
fi
: "${NGC_CLI_API_KEY:?Set NGC_API_KEY, NVIDIA_API_KEY, or NGC_CLI_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_CLI_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm -it --name molmim \
--runtime=nvidia \
-e CUDA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
-e NGC_CLI_API_KEY \
-v "${LOCAL_NIM_CACHE}:/home/nvs/.cache/nim" \
-p 8000:8000 \
nvcr.io/nim/nvidia/molmim:1.0.0
```
Readiness check:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
Local embedding smoke test after readiness. Local inference uses no
`Authorization` header:
```python
import requests
seed = "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
response = requests.post(
"http://localhost:8000/embedding",
headers={"Content-Type": "application/json"},
json={"sequences": [seed]},
timeout=60,
)
response.raise_for_status()
embedding_data = response.json()
embeddings = embedding_data["embeddings"]
print(f"received {len(embeddings)} embedding vector(s)")
```
## Hosted Generation Pattern
Use hosted `/generate` for seed-SMILES generation or optimization. Use
`algorithm: "CMA-ES"` for guided property optimization and `algorithm: "none"`
for unguided sampling around the seed.
```python
import os
import requests
hosted = True
url = (
"https://health.api.nvidia.com/v1/biology/nvidia/molmim/generate"
if hosted else "http://localhost:8000/generate"
)
headers = {"Content-Type": "application/json"}
if hosted:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"smi": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C",
"algorithm": "CMA-ES",
"num_molecules": 10,
"property_name": "QED",
"minimize": False,
"min_similarity": 0.4,
"particles": 8,
"iterations": 3,
}
response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
```
Generation gotchas:
- Field name is `smi`, not `smiles`.
- `algorithm` is `"CMA-ES"` or `"none"`.
- `property_name` is `"QED"` or `"plogP"`.
- `num_molecules` is 1-100. `iterations` is 1-1000. `particles` is 2-1000.
- `min_similarity` is 0-1 in the hosted API reference; local docs emphasize
common values up to 0.7 for constrained optimization.
- `scaled_radius` is 0-2 and is mainly used with `algorithm: "none"` or local
`/sampling`.
## Local Latent Workflow
Use local-only endpoints for embedding, hidden-state manipulation, and decode.
This is also the surface used by the guided optimization example package.
For local latent workflows, state explicitly that the hosted API reference
exposes `/generate`; `/embedding`, `/hidden`, `/decode`, and `/sampling` are
local-only in the current docs.
```python
seed = "CC(Cc1ccc(cc1)C(C(=O)O)C)C"
base = "http://localhost:8000"
headers = {"Content-Type": "application/json"}
embedding = requests.post(
f"{base}/embedding",
headers=headers,
json={"sequences": [seed]},
timeout=60,
)
embedding.raise_for_status()
embedding_data = embedding.json()
embeddings = embedding_data["embeddings"]
print(f"received {len(embeddings)} embedding vector(s)")
hidden = requests.post(
f"{base}/hidden",
headers=headers,
json={"sequences": [seed]},
timeout=60,
)
hidden.raise_for_status()
hidden_data = hidden.json()
hiddens = hidden_data["hiddens"]
mask = hidden_data["mask"]
decoded = requests.post(
f"{base}/decode",
headers=headers,
json={"hiddens": hiddens, "mask": mask},
timeout=60,
)
decoded.raise_for_status()
sampled = requests.post(
f"{base}/sampling",
headers=headers,
json={"sequences": [seed], "num_molecules": 10, "scaled_radius": 0.7},
timeout=60,
)
sampled.raise_for_status()
```
## Save And Validate Output
Save generated SMILES and validate before using them downstream.
```python
from pathlib import Path
import json
def molmim_smiles(result):
values = []
if isinstance(result.get("generated"), list):
for item in result["generated"]:
if isinstance(item, str):
values.append(item)
elif isinstance(item, list):
values.extend(x for x in item if isinstance(x, str))
molecules = result.get("molecules")
if isinstance(molecules, str):
molecules = json.loads(molecules)
if isinstance(molecules, list):
for item in molecules:
if isinstance(item, dict) and isinstance(item.get("sample"), str):
values.append(item["sample"])
return values
generated = molmim_smiles(result)
if not generated:
raise RuntimeError(f"MolMIM returned no generated molecules: {result}")
Path("molmim_response.json").write_text(json.dumps(result, indent=2))
Path("molmim_generated.smi").write_text("\n".join(generated) + "\n")
for i, smiles in enumerate(generated, start=1):
print(i, smiles)
```
Use RDKit when available to check parseability, uniqueness, simple property
ranges, and whether seed similarity constraints are plausible. Generated
molecules are candidates, not validated hits; use downstream property, docking,
affinity, toxicity, and synthetic-feasibility checks before prioritization.
## Troubleshooting
- Hosted `404` on `/embedding`, `/hidden`, `/decode`, or `/sampling`: those
endpoints are local-only in the docs.
- `401`: missing or unauthorized NGC key for hosted requests.
- Hosted response parsing: live hosted `/generate` may return `molecules` as
a JSON string of `{sample, score}` objects, while local endpoints may return
`generated`; parse both.
- `422`: invalid SMILES, unsupported `algorithm`, invalid `property_name`, or
parameter outside documented ranges.
- Local startup auth: set `NGC_CLI_API_KEY`, or set `NGC_API_KEY`/`NVIDIA_API_KEY`
and map it as shown above.
- Local startup cache misses: mount `LOCAL_NIM_CACHE` to `/home/nvs/.cache/nim`,
not `/opt/nim/.cache`.
Referenced files: 5
msa-search-nim16.4 KB
---
name: msa-search-nim
description: >
Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local Docker deployment. For local deployment, download the databases in parallel with aria2c and launch via NIM_MODEL_NAME (the recommended default fast path, ~14 min vs >80 min for the built-in downloader); a plain docker run uses the slow built-in downloader.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# MSA-Search NIM
Generate protein MSAs with GPU-accelerated MMSeqs2. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: MSA purpose, pairing/templates, limits, handoffs.
- `references/parameters.md`: database, pairing, depth, and template tuning.
- `references/validation.md`: alignment, template, and artifact checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode And Endpoint
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted standard MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict`
- Hosted paired MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict`
- Local standard MSA: `http://localhost:8000/biology/colabfold/msa-search/predict`
- Local paired MSA: `http://localhost:8000/biology/colabfold/msa-search/paired/predict`
- Local templates: `http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict`
Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
The hosted template path returned HTTP 404 in validation, so use local Docker
for template search unless the hosted docs/service changes.
## Local Docker
**Default local deployment = parallel download + `NIM_MODEL_NAME`.** The first recipe below
is the one to use for real workflows. It downloads the database(s) with a range-parallel
downloader (aria2c) and starts the NIM against those files — ~14 min for UniRef30 vs >80 min
for the NIM's built-in downloader (measured, H100). Do **not** reach for the plain `docker run`
(the "Fallback" subsection) unless you only want a `databases:pdb70` smoke test or you
deliberately want the NIM to manage its own blob cache.
Local setup requires a GPU. Size the NVMe volume to the profile you pick (UniRef30 ~490 GB;
full set ~1.4 TB). For setup answers, include env preflight, `docker login`, the parallel
download, `NIM_MODEL_NAME` launch, readiness, and then no-auth local inference. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.
```bash
# --- env preflight (do not drop the NVIDIA_API_KEY fallback) ---
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${DB_DIR:=/data/fast-db}" # where the parallel download lands
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
# --- 1) pick the DB version(s) you need (paired/complex work = uniref30 only) ---
DB_VERSION=uniref30_2302-m18v1
command -v aria2c >/dev/null || sudo apt-get install -y aria2
mkdir -p "$DB_DIR"
# --- 2) parallel download from NGC (see "Parallel Download" section for the all-DB loop) ---
curl -fsS -H "Authorization: Bearer $NGC_API_KEY" \
"https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/${DB_VERSION}/files" \
-o /tmp/files.json
DB_DIR="$DB_DIR" python3 - <<'PY'
import json, os
d = json.load(open("/tmp/files.json")); dbdir = os.environ["DB_DIR"]; lines = []
for url, path in zip(d["urls"], d["filepath"]):
lines += [url.strip(), f" dir={dbdir}", f" out={path}"]
open("/tmp/aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i /tmp/aria.in --max-concurrent-downloads=4 --max-connection-per-server=16 \
--split=16 --min-split-size=1M --continue=true --file-allocation=none
# --- 3) launch the NIM against the downloaded files (skips the slow built-in download) ---
docker run -d --name msa-search --runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_NAME=/databases \
-v "${DB_DIR}:/databases" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
If the DB is already present in `$DB_DIR`, skip steps 1-2 — the launch alone is a ~20 s warm
start. See "Parallel Download For Any Database Set" for the multi-database (`databases:all`)
loop and the full rationale.
### Fallback: Let The NIM Download Its Own Databases (slower)
Use this only for a quick `databases:pdb70` smoke test, or when you specifically want the NIM
to manage its own blob cache. It uses the built-in downloader, which is slow on large profiles
(UniRef30 stalled past 80 min in testing). Pin the smallest profile with `NIM_MODEL_PROFILE`
(see "Faster Startup") so it does not fetch the full 1.4 TB.
```bash
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
mkdir -p "${LOCAL_NIM_CACHE}"; chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm --name msa-search \
--runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2
```
## Faster Startup: Task-Specific Database Profiles
The full database download is ~1.4 TB and can take well over an hour on first launch. If you
only need some databases, select a **task-specific profile** so the NIM downloads just those.
This is the single biggest lever on local startup time.
List the profiles your image actually ships (hashes change between releases — never hardcode
them):
```bash
docker run --rm --entrypoint list-model-profiles nvcr.io/nim/colabfold/msa-search:2
```
Then pass the chosen hash with `NIM_MODEL_PROFILE`:
```bash
docker run --rm --name msa-search \
--runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2
```
Profiles available in this image (confirm hashes with `list-model-profiles`):
| Profile tags | Databases | Best for | Storage |
|---|---|---|---|
| `databases:pdb70` | PDB70 | Quick testing / smoke check | ~100 MB |
| `databases:uniref30` | UniRef30 | **Paired MSA search for complexes** — UniRef30 is the only DB used for species-based pairing | ~500 GB |
| `databases:uniref30,pdb70,pdb` | UniRef30 + PDB70 + PDB structures | Structural template search | ~700 GB |
| `databases:all` (default) | UniRef30 + ColabFold envdb + PDB70 + PDB100 + PDB structures | Full sensitivity, all databases | ~1.2 TB |
Verify the loaded profile after readiness:
```bash
curl -s localhost:8000/v1/metadata | jq
```
Notes:
- The request-level `databases` parameter only selects among databases **already
downloaded**; it does NOT change what is fetched at startup. Startup footprint is set by
`NIM_MODEL_PROFILE` alone.
- **Paired search needs UniRef30 only.** `colabfold_envdb_202108` has no taxonomy and cannot
be used for pairing, so `databases:uniref30` is the correct, smallest profile for
complex/paired workflows — it skips the envdb, the largest part of the full set.
- For maximum monomer sensitivity (UniRef30 + envdb merged) you still need `databases:all`;
there is no envdb-inclusive profile smaller than the full set.
### Custom Or Individual Databases
To use a single manually downloaded database (or your own MMSeqs2 DB), download it from NGC
and point the NIM at the mount with `NIM_MODEL_NAME` instead of a profile:
```bash
ngc registry model download-version nim/colabfold/msa-search:uniref30_2302-m18v1
# then mount the directory and set -e NIM_MODEL_NAME=/databases
```
`NIM_MODEL_NAME` **replaces** the profile databases entirely — the NIM uses only what is
under that directory (discovered by scanning for `**/*.idx`). Mount multiple databases under
one parent to combine them. NGC-downloaded databases are pre-indexed for GPU Server; custom
databases must be indexed with `mmseqs createindex` first. Individually downloadable NGC model
versions: `uniref30_2302-m18v1`, `colabfold_envdb_202108-m18v1`, `pdb70_220313-m18v1`,
`pdb100_230517-m18v1`, `pdb_20251028_zip-m18v1`.
## Recommended: Parallel Download For Any Database Set (Fast Deployment)
**This is the recommended way to download the databases at all — for any profile,
including the full `databases:all` set.** Task-specific profiles cut *what* you
download; this parallel downloader cuts *how long* that download takes. Use it whether
you need one database or all of them. The gain is largest for the `databases:uniref30` profile, which is ~490 GB
dominated by two very large files (a ~241 GB GPU index and a ~134 GB sequence DB).
The NIM's built-in downloader parallelizes **across files** (`max_parallel_files=10`) but pulls
each file over roughly one connection. The NGC CDN throttles a single connection to ~20–25
MB/s, so while the downloader is fetching one of the two giant files, most of its parallel
slots sit idle and throughput collapses to that single-flow rate. Measured on an H100 node,
the built-in path did not reach `/health/ready` in over 80 minutes.
A range-parallel downloader splits **each file** into many byte-range segments (the NGC CDN
advertises `accept-ranges: bytes`), so a single 241 GB file is pulled over 16 connections at
once — ~15× the single-flow rate. Same node, `aria2c` fetched the full ~490 GB in **~13.5
minutes**.
Workflow (download once with aria2, then start the NIM against the files via `NIM_MODEL_NAME`):
```bash
# 1) Get presigned file URLs for the individual database model version from NGC.
# (Requires NGC_API_KEY. The response arrays `urls` and `filepath` are positionally paired.)
curl -s -H "Authorization: Bearer $NGC_API_KEY" \
'https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/uniref30_2302-m18v1/files' \
-o files.json
# 2) Build an aria2 input file (URL + target filename per entry) and download in parallel.
python3 - <<'PY'
import json
d = json.load(open("files.json"))
lines = []
for url, path in zip(d["urls"], d["filepath"]):
lines += [url.strip(), " dir=/data/fast-db", f" out={path}"]
open("aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i aria.in \
--max-concurrent-downloads=4 --max-connection-per-server=16 --split=16 \
--min-split-size=1M --continue=true --file-allocation=none
# 3) Start the NIM against the downloaded directory. NIM_MODEL_NAME makes the NIM discover
# databases by scanning for **/*.idx, bypassing the profile/blob cache entirely.
docker run -d --name msa-search --runtime=nvidia --gpus all \
-e NGC_API_KEY \
-e NIM_MODEL_NAME=/databases \
-v /data/fast-db:/databases \
-p 8000:8000 \
nvcr.io/nim/colabfold/msa-search:2
```
For **all databases** (equivalent to `databases:all`), repeat step 1 for each individual DB
version and download them into sibling directories under one parent, then point
`NIM_MODEL_NAME` at that parent — the NIM discovers every DB by scanning `**/*.idx`:
```bash
# fetch each DB's file list into /data/all-db/<db>/ ... then one aria2c per list, e.g.:
for V in uniref30_2302-m18v1 colabfold_envdb_202108-m18v1 pdb70_220313-m18v1 \
pdb100_230517-m18v1 pdb_20251028_zip-m18v1; do
curl -s -H "Authorization: Bearer $NGC_API_KEY" \
"https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/$V/files" \
-o "files_$V.json"
# build an aria2 input from files_$V.json (dir=/data/all-db) and run aria2c on it
done
# then launch once against the parent:
# docker run -d ... -e NIM_MODEL_NAME=/databases -v /data/all-db:/databases ...
```
The per-connection CDN throttle is the same for every database, so parallel download helps
the full set proportionally — the more you download, the more absolute time it saves.
Notes:
- The presigned URLs expire (typically within a day) — build the aria2 input and start the
download promptly after fetching `files.json`.
- Keep the downloaded directory's internal layout intact (e.g. `uniref30_2302/…`); the
`filepath` values already encode it. The NIM needs the `.idx` file plus its companion files
and the small `.UNIREF30_READY` / `*.tar.gz.unpacked` markers.
- The bottleneck is the CDN's per-connection cap, not local disk or CPU — a fast NVMe volume
writes far faster than the network delivers. Raising `--split` / `--max-connection-per-server`
helps only up to the node's aggregate egress ceiling.
- Best of all: download the profile once, then **persist the cache volume** (or this
`fast-db` directory) and mount it on future nodes for a ~20 s warm start with no re-download.
## Standard MSA Request
Use exact case-sensitive database names and response keys.
```python
import os
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict"
if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"sequence": "SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR",
"databases": ["Uniref30_2302", "colabfold_envdb_202108"],
"e_value": 0.0001,
"output_alignment_formats": ["a3m"],
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
## Paired MSA Request
Use paired search for protein complexes; payload field is `sequences` plural,
and output is `alignments_by_chain`.
```python
url = (
"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict"
if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/paired/predict"
)
payload = {
"sequences": [chain_a_sequence, chain_b_sequence],
"e_value": 0.0001,
"output_alignment_formats": ["a3m"],
}
```
## Local Template Search
Use local Docker for structural templates. Set `max_msa_sequences=500` unless
`NIM_GLOBAL_MAX_MSA_DEPTH` was changed.
```python
url = "http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict"
headers = {"Content-Type": "application/json"}
payload = {
"sequence": "VLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHGKKVADALTNAVA",
"structural_template_databases": ["pdb70_220313"],
"max_structures": 20,
"max_msa_sequences": 500,
}
```
## Save Outputs
```python
# Standard MSA: result["alignments"][database][format]["alignment"]
for db_name, formats in result.get("alignments", {}).items():
for fmt_name, data in formats.items():
with open(f"msa_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
handle.write(data["alignment"])
# Paired MSA: one alignment set per chain
for chain_id, chain_data in result.get("alignments_by_chain", {}).items():
for db_name, formats in chain_data.items():
for fmt_name, data in formats.items():
with open(f"msa_chain_{chain_id}_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
handle.write(data["alignment"])
# Template search: save mmCIF structures and M8 hit tables
for name, cif in result.get("structures", {}).items():
open(f"template_{name}.cif", "w", encoding="utf-8").write(cif)
for name, hit_table in result.get("search_hits", {}).items():
open(f"template_hits_{name}.m8", "w", encoding="utf-8").write(hit_table)
```
A3M output can feed OpenFold3, AlphaFold2, or RoseTTAFold. For alignment depth,
template, and sequence sanity checks, read `references/validation.md`.
## Limits And Troubleshooting
- Sequence length: 1-4096 amino acids; `X` works since v2.3.0.
- `max_msa_sequences`: 1-500; local GPU server default must match
`NIM_GLOBAL_MAX_MSA_DEPTH`.
- Paired MSA requires at least two sequences.
- Local URL 404 usually means an accidental `/v1/` prefix.
- First local run can take hours while databases populate `LOCAL_NIM_CACHE`.
Referenced files: 5
msa-structure-prediction-pipeline6.2 KB
---
name: msa-structure-prediction-pipeline
description: >
Run a complete protein structure prediction pipeline using NVIDIA BioNeMo NIMs:
search for MSA alignments with MSA-Search (ColabFold), then predict the structure
with OpenFold3 using the retrieved alignments. Use this skill whenever the user wants
to predict a protein structure with maximum accuracy using MSA context, run the
full AlphaFold3-style pipeline, generate MSA-informed structure predictions, or
improve structure prediction accuracy by providing evolutionary information.
Triggers on: MSA structure prediction pipeline, structure prediction pipeline, MSA-informed prediction, OpenFold3,
ColabFold MSA, AlphaFold3 pipeline, protein structure, homology search, a3m alignment,
UniRef30, NIM microservice. This pipeline chains MSA-Search and OpenFold3.
license: Apache-2.0 AND CC-BY-4.0
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# MSA Structure Prediction Pipeline
Predict protein structures with high accuracy by chaining two BioNeMo NIMs:
```
Step 1: MSA-Search → Step 2: OpenFold3
(Search homologs) (Predict structure with MSA)
```
---
## Overview
Why chain these NIMs?
- **MSA-Search** finds evolutionary homologs in UniRef30 and ColabFold databases using GPU-accelerated MMSeqs2. The resulting alignment provides crucial evolutionary information.
- **OpenFold3** uses the MSA to improve structure prediction accuracy — especially for sequences where no close homolog exists in PDB.
- Running MSA-Search first means OpenFold3 gets the full evolutionary context rather than a single-sequence prediction.
---
## Before you start
Confirm with the user:
1. **Query sequence**: amino acid sequence to predict
2. **MSA depth**: how many sequences to retrieve (default 500; more = slower but more context)
3. **API mode**: hosted or local Docker?
Note: local MSA-Search requires 1.4 TB of database storage — strongly recommend hosted unless the user has that infrastructure.
For local Docker, do not assume MSA-Search and OpenFold3 are both on
`localhost:8000` concurrently. Run one container at a time and hand off the A3M
file, or start each NIM on a distinct host port and set the URLs explicitly.
---
## Step 1: Search for MSA with MSA-Search
```python
import requests, json, os
from pathlib import Path
NGC_API_KEY = os.environ["NGC_API_KEY"]
HOSTED = True
query_sequence = "<YOUR_PROTEIN_SEQUENCE>"
if HOSTED:
msa_url = "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict"
headers = {"Content-Type": "application/json",
"Authorization": f"Bearer {NGC_API_KEY}"}
else:
msa_url = "http://localhost:8000/biology/colabfold/msa-search/predict"
headers = {"Content-Type": "application/json"}
payload = {
"sequence": query_sequence,
"databases": ["Uniref30_2302", "colabfold_envdb_202108"],
"e_value": 0.0001,
"output_alignment_formats": ["a3m"],
}
r = requests.post(msa_url, headers=headers, json=payload)
r.raise_for_status()
msa_result = r.json()
# Extract the A3M alignment
a3m_alignment = msa_result["alignments"]["Uniref30_2302"]["a3m"]["alignment"]
# Save for reference
with open("query_msa.a3m", "w") as f:
f.write(a3m_alignment)
# Count sequences in alignment
n_seqs = a3m_alignment.count(">")
print(f"Step 1 complete: found {n_seqs} homologous sequences")
print(f"MSA saved to query_msa.a3m")
```
---
## Step 2: Predict structure with OpenFold3
Pass the MSA directly into OpenFold3's `msa` field:
```python
if HOSTED:
of3_url = "https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict"
else:
of3_url = "http://localhost:8000/biology/openfold/openfold3/predict"
# Build the OpenFold3 MSA structure from the retrieved alignment
msa_data = {
"uniref30": {
"a3m": {
"alignment": a3m_alignment,
"format": "a3m"
}
}
}
# Optionally also include colabfold_envdb alignment if requested
# env_alignment = msa_result["alignments"]["colabfold_envdb"]["a3m"]["alignment"]
# msa_data["colabfold_env"] = {"a3m": {"alignment": env_alignment, "format": "a3m"}}
payload = {
"inputs": [{
"input_id": "prediction_with_msa",
"output_format": "pdb",
"molecules": [
{
"type": "protein",
"sequence": query_sequence,
"diffusion_samples": 1,
"msa": msa_data
}
]
}]
}
r = requests.post(of3_url, headers=headers, json=payload, timeout=300)
r.raise_for_status()
result = r.json()
output = result["outputs"][0]
for i, sample in enumerate(output["structures_with_scores"]):
fmt = sample["format"]
filename = f"predicted_structure_{i+1}.{fmt}"
with open(filename, "w") as f:
f.write(sample["structure"])
print(f"\nStep 2 complete: {filename} saved")
print(f" Confidence: {sample['confidence_score']:.4f}")
print(f" pLDDT: {sample['complex_plddt_score']:.4f}")
print(f" pTM: {sample['ptm_score']:.4f}")
```
---
## Comparing single-sequence vs MSA-informed prediction
If the user wants to see the impact of MSA, run OpenFold3 twice — once with the full MSA and once with just the query sequence as a minimal alignment:
```python
# Minimal MSA (single sequence — same as no MSA context):
minimal_msa = {
"main": {
"a3m": {
"alignment": f">query\n{query_sequence}",
"format": "a3m"
}
}
}
```
A larger, higher-quality MSA typically yields higher pLDDT and lower pDE, especially for proteins with many known homologs.
---
## For protein complexes
Use the `/paired/predict` endpoint of MSA-Search to get paired alignments for multi-chain complexes, then pass each chain's alignment into the corresponding molecule's `msa` field and `paired_msa` fields:
```python
# Paired MSA search endpoint for complexes:
msa_paired_url = "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict"
paired_payload = {
"sequences": [chain_A_sequence, chain_B_sequence],
"e_value": 0.0001,
}
```
---
## Quick reference — skill dependencies
| Step | Skill | Key endpoint |
|---|---|---|
| MSA search | `msa-search-nim` | `/biology/colabfold/msa-search/predict` |
| Structure prediction | `openfold3-nim` | `/biology/openfold/openfold3/predict` |
nvmolkit-usage17.3 KB
---
name: nvmolkit-usage
description: >-
Write code that calls the installed nvMolKit Python API for GPU-accelerated,
batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine
similarity, ETKDG conformer embedding, MMFF/UFF optimization, TFD, conformer
RMSD, Butina clustering, and substructure search. Use when the user is
importing `nvmolkit.*`, debugging an `nvmolkit` call, choosing between
nvMolKit and RDKit for a batched cheminformatics workflow, or wiring nvMolKit
results into a torch/numpy pipeline. Out of scope: building nvMolKit from
source.
license: Apache-2.0
metadata:
owner: Kevin Boyd (@scal444)
risk_tier: skill
---
# nvMolKit usage
## What nvMolKit is
GPU-accelerated, batched implementations of common RDKit operations. APIs mirror RDKit where possible but are batch-oriented: they take lists of `rdkit.Chem.Mol` (or lists of fingerprints) and process them in parallel on one or more GPUs. nvMolKit links against RDKit at build time; inputs and outputs are real RDKit `Mol` objects.
## Where nvMolKit does well
Reach for nvMolKit when:
- The workload is **a large batch of molecules** processed together (typically thousands or more).
- The metric is **throughput / total wall time across the batch**, not per-molecule latency.
- The same operation is **repeated identically** across the batch (fingerprinting a library, embedding/minimizing many conformers, bulk pairwise similarity), so the GPU stays saturated.
Plain RDKit is usually the better choice for single-molecule one-offs or workflows that can't be expressed as a batch. nvMolKit is not meant to replace RDKit for those cases.
## Runtime requirements
- An NVIDIA GPU with compute capability 7.0 (V100) or higher
- A CUDA driver compatible with CUDA 12.6+.
- A working `torch` install with CUDA support (nvMolKit returns GPU tensors via `torch`'s CUDA array interface).
If CUDA is unavailable, nvMolKit calls raise. There is no CPU fallback - if the user needs one, use RDKit directly for that path.
When helping with installation, make the user choose a PyTorch CUDA backend that the host driver supports before installing nvMolKit. nvMolKit's PyPI wheels are built with CUDA Toolkit 12.9 and depend on CUDA 12 runtime packages, but pip/uv can still select a CUDA 13 PyTorch wheel unless the install command says otherwise.
- Conda: prefer conda-forge `pytorch-gpu`; pin `cuda-version=12.6` or another CUDA version supported by the driver.
- pip: send the user to the PyTorch install selector (`https://pytorch.org/get-started/locally/`) or previous-versions page (`https://pytorch.org/get-started/previous-versions/`) to install `torch` for a CUDA 12.x backend before installing nvMolKit.
- uv: install nvMolKit with an explicit backend, e.g. `uv pip install --torch-backend=cu128 nvmolkit`.
## Verify the install before writing real code
Run this once to confirm nvMolKit is importable and a GPU op works end to end:
```python
import nvmolkit
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
print("nvmolkit:", nvmolkit.__version__)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
result = fpgen.GetFingerprints(mols)
torch.cuda.synchronize()
fps = result.torch()
print("fps shape:", tuple(fps.shape), "dtype:", fps.dtype)
# Expected: shape (3, 32), dtype torch.int32 (1024 bits packed into 32 int32s per row)
```
If this fails, point the user at the install guide on the docs site rather than guessing - see "Going deeper" below.
## Entry points
| Task | Module | Primary entry point |
|---|---|---|
| Morgan fingerprints | `nvmolkit.fingerprints` | `MorganFingerprintGenerator(radius, fpSize).GetFingerprints(mols)` |
| Bulk Tanimoto / cosine similarity | `nvmolkit.similarity` | `crossTanimotoSimilarity(...)`, `crossCosineSimilarity(...)`, plus `*MemoryConstrained` variants for results too large to fit in GPU memory |
| ETKDG conformer embedding | `nvmolkit.embedMolecules` | `EmbedMolecules(molecules, params, confsPerMolecule, ...)` |
| MMFF94 optimization (one-shot) | `nvmolkit.mmffOptimization` | `MMFFOptimizeMoleculesConfs(molecules, ...)` |
| UFF optimization (one-shot) | `nvmolkit.uffOptimization` | `UFFOptimizeMoleculesConfs(molecules, ...)` |
| Forcefield with custom options + constraints | `nvmolkit.batchedForcefield` | `MMFFBatchedForcefield(mols, properties=..., nonBondedThreshold=..., ignoreInterfragInteractions=..., hardwareOptions=...)`, `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`. Per-molecule view `ff[i]` exposes `add_distance_constraint`, `add_position_constraint`, `add_angle_constraint`, `add_torsion_constraint`. Methods: `.compute_energy()`, `.compute_gradients()`, `.minimize(maxIters, forceTol)` |
| Pairwise conformer RMSD | `nvmolkit.conformerRmsd` | `GetConformerRMSMatrix(mol)`, `GetConformerRMSMatrixBatch(mols)` |
| Torsion Fingerprint Deviation (TFD) | `nvmolkit.tfd` | `GetTFDMatrix(mol)`, `GetTFDMatrices(mols)` |
| Butina clustering | `nvmolkit.clustering` | `butina(distance_matrix, cutoff)` (precomputed matrix), `fused_butina(fingerprints, cutoff)` (memory-efficient, on-the-fly) |
| Substructure search | `nvmolkit.substructure` | `hasSubstructMatch`, `countSubstructMatches`, `getSubstructMatches` |
| Hardware tuning (batch size, GPU IDs) | `nvmolkit.types` | `HardwareOptions(...)` passed to ETKDG / MMFF / UFF |
| Optional autotuning of `HardwareOptions` | `nvmolkit.autotune` | `tune_embed_molecules`, `tune_mmff_optimize`, `tune_uff_optimize`, `tune_batched_forcefield`. Requires the `optuna` package |
## Result types and execution model
Two return shapes carry GPU-resident output, depending on what the operation produces.
### `AsyncGpuResult`
Used by operations that return a single flat tensor (fingerprints, similarity matrices, RMSD/TFD vectors, Butina inputs). Key behaviors:
- Asynchronous. The kernel may not have completed when the call returns.
- `result.torch()` returns a zero-copy `torch.Tensor` on the GPU. Caller is responsible for synchronizing before reading values on the host.
- `result.numpy()` synchronizes and returns a CPU numpy array.
- Exposes `__cuda_array_interface__`, so it can be passed directly into other nvMolKit functions (e.g. fingerprints → similarity) with no host round-trip.
#### CUDA stream control
A subset of the `AsyncGpuResult`-returning APIs accept an optional `stream: torch.cuda.Stream | None = None` argument so callers can submit nvMolKit work to a non-default stream and overlap it with their own kernels. When omitted, the call uses the current torch stream.
APIs that take a `stream` argument:
- `MorganFingerprintGenerator.GetFingerprints`
- `crossTanimotoSimilarity`, `crossCosineSimilarity`, and their `*MemoryConstrained` variants
- `butina`, `fused_butina`
- `GetConformerRMSMatrix`, `GetConformerRMSMatrixBatch`
Other APIs (ETKDG, MMFF/UFF optimization, TFD, substructure search) are synchronous to the caller — no stream plumbing needed.
Typical pattern:
```python
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity
stream = torch.cuda.Stream()
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
with torch.cuda.stream(stream):
fps = fpgen.GetFingerprints(mols, stream=stream)
sim = crossTanimotoSimilarity(fps, stream=stream)
stream.synchronize()
print(sim.torch())
```
### `Device3DResult`
Used by ETKDG embedding and MMFF/UFF optimization (one-shot and `BatchedForcefield`) when called with `output=CoordinateOutput.DEVICE`. The GPU-resident equivalent of writing conformers back to `Mol` objects. Fields:
- `values`: `AsyncGpuResult` of shape `(total_atoms, 3)` float64. Concatenated conformer coordinates in CSR-style layout.
- `atom_starts`, `mol_indices`, `conf_indices`: `AsyncGpuResult` int32 buffers describing the layout (`values[atom_starts[i]:atom_starts[i+1]]` is conformer `i`'s atoms).
- `energies`, `converged`: `AsyncGpuResult` buffers populated only for MMFF/UFF minimization (not for plain ETKDG).
- `gpu_id`: device the buffers live on. The `targetGpu` argument on each API picks this; `targetGpu=-1` uses the default consolidation device.
- `.per_molecule()` returns nested `list[list[torch.Tensor]]` of per-conformer views; `.dense(pad_value=nan)` materializes a padded `(n_mols, max_confs, max_atoms, 3)` tensor.
The default mode (`CoordinateOutput.RDKIT_CONFORMERS`) still writes optimized coordinates back into each `Mol` and returns Python lists of energies/convergence flags. Reach for `CoordinateOutput.DEVICE` when chaining downstream GPU work (e.g. ETKDG → MMFF → similarity scoring) without host round-trips.
## Configuration
Two configuration objects expose the GPU/CPU knobs.
### `HardwareOptions` (ETKDG, MMFF, UFF)
`from nvmolkit.types import HardwareOptions`. Passed via `hardwareOptions=` to `EmbedMolecules`, `MMFFOptimizeMoleculesConfs`, `UFFOptimizeMoleculesConfs`, and the `BatchedForcefield` constructors. Every field has an "auto" sentinel; the defaults are usually fine.
| Field | Type | Default | Meaning |
|---|---|---|---|
| `preprocessingThreads` | int | `-1` (all visible CPUs) | CPU threads for preprocessing |
| `batchSize` | int | `-1` (auto-tuned) | Number of conformers per GPU batch |
| `batchesPerGpu` | int | `-1` (auto) | Concurrent batches per GPU; must be `>0` or `-1` |
| `gpuIds` | `list[int]` | `[]` (all visible GPUs) | Specific device ordinals to target |
Passing a `gpuIds` entry for a device that isn't visible raises `RuntimeError: invalid device ordinal`. For finding good values automatically across a representative sample, see `nvmolkit.autotune` (requires the `optuna` extra); each `tune_*` function returns a `TuneResult` whose `best_config` is a fully-populated `HardwareOptions` ready to pass back into the real call.
`HardwareOptions` round-trips through `to_dict()` / `from_dict()` for persisting tuned configs to disk.
### `SubstructSearchConfig` (substructure search)
`from nvmolkit.substructure import SubstructSearchConfig`. Passed via `config=` to `hasSubstructMatch`, `countSubstructMatches`, and `getSubstructMatches`.
| Field | Type | Default | Meaning |
|---|---|---|---|
| `batchSize` | int | `1024` | (target, query) pairs per GPU batch |
| `workerThreads` | int | `-1` (auto) | GPU runner threads per GPU |
| `preprocessingThreads` | int | `-1` (auto) | CPU threads for preprocessing |
| `maxMatches` | int | `0` (unlimited) | Max matches returned per (target, query) pair |
| `uniquify` | bool | `False` | Drop duplicate matches that differ only in atom enumeration order |
| `gpuIds` | `list[int] \| None` | `None` (current device only) | Specific device ordinals to target |
Substructure search currently does not support chirality-aware matching, enhanced stereochemistry, or other advanced RDKit `SubstructMatchParameters` options.
## Recipes
### Morgan fingerprints + bulk Tanimoto similarity
```python
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity
smiles = ["CCO", "CCN", "c1ccccc1", "CC(=O)O", "CCOCC"]
mols = [Chem.MolFromSmiles(smi) for smi in smiles]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
fps = fpgen.GetFingerprints(mols)
sim = crossTanimotoSimilarity(fps)
torch.cuda.synchronize()
print(sim.torch())
```
Inputs are `list[Mol]`. Output of `GetFingerprints` is an `AsyncGpuResult` wrapping an `(n_mols, fpSize / 32)` int32 tensor of packed bits. Pass it straight into `crossTanimotoSimilarity` for an `(n, n)` similarity matrix; pass two fingerprint sets for an `(n, m)` cross-matrix. For sets too large to materialize on the GPU, use `crossTanimotoSimilarityMemoryConstrained` (chunked compute, returns numpy on CPU).
### ETKDG conformer embedding
```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import ETKDGv3
from nvmolkit.embedMolecules import EmbedMolecules
mols = [AddHs(MolFromSmiles(smi)) for smi in ["C1CCCCC1", "C1CCCCC2CCCCC12", "COO"]]
params = ETKDGv3()
params.useRandomCoords = True
EmbedMolecules(mols, params, confsPerMolecule=10, maxIterations=-1)
for mol in mols:
print(mol.GetNumConformers())
```
Inputs are `list[Mol]`, sanitized and with hydrogens added (`AddHs`). Conformers are added in-place. `params.useRandomCoords` must be `True` - nvMolKit's ETKDG only supports random-coord initialization. A handful of niche `EmbedParameters` options are not supported (bounds matrices, custom CPCI, coord maps, separate-fragment embedding); the Features section of the docs site lists the full restrictions.
### MMFF94 minimization of a batch of conformers
```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import ETKDGv3
from nvmolkit.embedMolecules import EmbedMolecules
from nvmolkit.mmffOptimization import MMFFOptimizeMoleculesConfs
mols = [AddHs(MolFromSmiles(smi)) for smi in ["CCO", "CCN", "c1ccccc1"]]
params = ETKDGv3(); params.useRandomCoords = True
EmbedMolecules(mols, params, confsPerMolecule=5)
energies = MMFFOptimizeMoleculesConfs(mols, maxIters=500)
for mol, mol_energies in zip(mols, energies):
print(mol.GetNumConformers(), mol_energies)
```
Inputs are `list[Mol]` with conformers already populated (typically by ETKDG, RDKit's `EmbedMultipleConfs`, or a prior nvMolKit call). Coordinates are updated in place; the return is `list[list[float]]` of optimized energies aligned with the input molecule order and conformer index. UFF is identical in shape: swap in `from nvmolkit.uffOptimization import UFFOptimizeMoleculesConfs`.
If any input molecule is `None` or lacks MMFF/UFF atom types, the call raises `ValueError`. The exception's `args[1]` is a dict with keys `"none"` and `"no_params"` listing the offending indices - useful for filtering a noisy input set.
### Conformer RMSD and Butina clustering
```python
import torch
from rdkit import Chem
from rdkit.Chem.rdDistGeom import EmbedMultipleConfs
from nvmolkit.clustering import butina
from nvmolkit.conformerRmsd import GetConformerRMSMatrixBatch
mols = [Chem.AddHs(Chem.MolFromSmiles(smi)) for smi in ["CCCCCC", "c1ccccc1"]]
for mol in mols:
EmbedMultipleConfs(mol, numConfs=10)
# Remove hydrogens after embedding for heavy-atom RMSD.
heavy_mols = [Chem.RemoveHs(mol) for mol in mols]
# Default RMSD output is RDKit-compatible condensed lower-triangle form.
condensed = GetConformerRMSMatrixBatch(heavy_mols)
# Butina expects a square distance matrix, so request square GPU tensors.
square = GetConformerRMSMatrixBatch(heavy_mols, output_format="square")
clusters = [butina(distance_matrix, cutoff=0.5).torch() for distance_matrix in square]
torch.cuda.synchronize()
for mol_clusters in clusters:
print(mol_clusters.cpu().tolist())
```
`GetConformerRMSMatrix(mol)` and `GetConformerRMSMatrixBatch(mols)` default to `output_format="condensed"`, returning `AsyncGpuResult` objects that wrap RDKit-style flat vectors of length `N * (N - 1) // 2`. Use `output_format="square"` when chaining into `butina()` or any other API that expects an `N x N` distance matrix. Both forms live on the GPU; call `.numpy()` on condensed results or synchronize before moving square tensors to the CPU.
### Custom forcefield options + constraints (`BatchedForcefield`)
Reach for `MMFFBatchedForcefield` / `UFFBatchedForcefield` instead of the one-shot `MMFFOptimizeMoleculesConfs` / `UFFOptimizeMoleculesConfs` when you need any of:
- Custom `maxIters` / `forceTol` per call
- Per-molecule `nonBondedThreshold` (MMFF) or `vdwThreshold` (UFF), or per-molecule `ignoreInterfragInteractions`
- Per-molecule `MMFFMolProperties` objects (e.g. MMFF94s vs MMFF94)
- Distance, position, angle, or torsion constraints
- Standalone `compute_energy()` / `compute_gradients()` without minimization
```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import EmbedMultipleConfs
from nvmolkit.batchedForcefield import MMFFBatchedForcefield
mols = [AddHs(MolFromSmiles(smi)) for smi in ["CCO", "CCCCCC"]]
for mol in mols:
EmbedMultipleConfs(mol, numConfs=5)
ff = MMFFBatchedForcefield(
mols,
nonBondedThreshold=[100.0, 20.0],
ignoreInterfragInteractions=True,
)
ff[0].add_position_constraint(0, max_displ=0.1, force_constant=50.0)
ff[1].add_distance_constraint(0, 4, relative=False, min_len=1.8, max_len=2.2, force_constant=25.0)
energies, converged = ff.minimize(maxIters=500, forceTol=1e-4)
for mol, mol_energies, mol_converged in zip(mols, energies, converged):
print(mol.GetNumConformers(), mol_energies, mol_converged)
```
All conformers of each input molecule are minimized in one batch. Constraints attached via `ff[i].add_*_constraint(...)` apply to every conformer of molecule `i`; constraint setters mark the wrapper dirty and the native forcefield rebuilds on the next call. Pass `output=CoordinateOutput.DEVICE` to `.minimize(...)` to keep optimized coordinates on the GPU (`Device3DResult`) instead of writing them back into RDKit conformers. UFF is the same shape: `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`.
## Going deeper
- Full feature list, API reference, and guides: <https://nvidia-bionemo.github.io/nvMolKit/>
- What changed in each release: <https://nvidia-bionemo.github.io/nvMolKit/changelog.html>
- Worked examples (Jupyter notebooks): the `examples/` directory in the GitHub repo
openfold2-nim7.49 KB
---
name: openfold2-nim
description: >
Use this skill for OpenFold2, NVIDIA's BioNeMo NIM microservice for monomer protein structure prediction. Invoke whenever the user mentions OpenFold2, AlphaFold2-like monomer folding, protein sequence-to-structure prediction, A3M MSAs, mmCIF templates, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# OpenFold2 NIM
Predict a single protein-chain structure from an amino-acid sequence, with
optional A3M multiple sequence alignments and mmCIF templates. Use this
`SKILL.md` for basic hosted/local NIM use; load supplemental files only when
the task needs deeper context:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: model scope, strengths, limitations, and handoffs.
- `references/parameters.md`: MSA, template, model-selection, and relax effects.
- `references/validation.md`: artifact and scientific sanity checks.
- `references/examples.md`: compact hosted/local payload patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted URL: `https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template`
- Local URL: `http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template`
- Local readiness: `http://localhost:8000/v1/health/ready`
Mode difference: hosted and local use the same prediction path except local
does not include `/v1/`. Hosted requests use `Authorization: Bearer
$NGC_API_KEY`; local inference requests use no auth header after readiness.
## Auth And Environment
Do not print API keys. Confirm they exist with shell tests, not echoes.
Hosted needs `NGC_API_KEY` in the request header. Supported local Docker
startup uses `NGC_API_KEY`, or `NVIDIA_API_KEY` as a fallback, plus
`LOCAL_NIM_CACHE`. A repo-root `.env` file may be sourced as a local override.
## Local Docker
Use the official OpenFold2 NIM image and mount `LOCAL_NIM_CACHE` at
`/opt/nim/.cache`. Current docs recommend at least 80 GB disk, 64 GB system
RAM, 8 CPU cores, and one supported GPU; the container is roughly 55 GB and
first startup downloads about 10 GB of model parameters.
When writing local setup commands, copy the preflight below exactly. Do not
drop `.env`, `NVIDIA_API_KEY`, `LOCAL_NIM_CACHE`, or the no-auth local request.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm --name openfold2 \
--runtime=nvidia \
--gpus "device=${NIM_TEST_GPU}" \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/openfold/openfold2:latest
```
Readiness check:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## Request Pattern
Use Python `requests`; curl escaping is fragile for A3M/mmCIF text. The
`sequence` field is required. `input_id`, `alignments`, `selected_models`,
`relax_prediction`, `use_templates`, and `explicit_templates` are optional.
```python
import os
import requests
hosted = True
url = (
"https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template"
if hosted
else "http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template"
)
headers = {"Content-Type": "application/json"}
if hosted:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
seq = "MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT"
payload = {
"sequence": seq,
"input_id": "kras_fragment",
"selected_models": [1],
"relax_prediction": False,
"alignments": {
"uniref90": {
"a3m": {
"alignment": f">query\n{seq}",
"format": "a3m",
}
}
},
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
Payload gotchas:
- OpenFold2 is monomer-only. For protein-ligand, protein-DNA/RNA, or
multi-chain complexes, use OpenFold3 or Boltz2 instead.
- `sequence` must use valid amino-acid IUPAC symbols.
- Hosted API docs list sequence length 1-1000; local docs say current NIM
supports sequences up to 2048 residues on supported hardware.
- A3M alignments go under `alignments` by database name, then `a3m` with
`alignment` and `format`. When the user needs to create or deepen an MSA,
hand off to `msa-search-nim` / MSA Search and map its A3M output into this
`alignments` shape.
- Starting with OpenFold2 2.0.0, use `explicit_templates` with mmCIF content;
do not write new HHR-template examples.
- `selected_models` chooses AlphaFold2/OpenFold parameter sets 1-5. Select one
or two models for smoke tests; use all five for stronger production runs.
## Save And Interpret Output
The response includes one prediction per selected model, ordered by confidence.
Save every returned structure-like text field and the full JSON response so
field-shape differences are auditable. Production answers should explicitly
write `.pdb` or `.cif` artifacts, preserve the response JSON, and print any
confidence/ranking fields the service returns.
```python
from pathlib import Path
import json
Path("openfold2_response.json").write_text(json.dumps(result, indent=2))
def save_strings(obj, prefix="openfold2"):
i = 0
if isinstance(obj, dict):
for key, value in obj.items():
if isinstance(value, str) and ("ATOM" in value or value.lstrip().startswith("data_")):
i += 1
ext = "cif" if value.lstrip().startswith("data_") else "pdb"
Path(f"{prefix}_{key}_{i}.{ext}").write_text(value)
elif isinstance(value, (dict, list)):
i += save_strings(value, f"{prefix}_{key}")
elif isinstance(obj, list):
for idx, value in enumerate(obj, start=1):
if isinstance(value, (dict, list)):
i += save_strings(value, f"{prefix}_{idx}")
return i
saved = save_strings(result)
print(f"saved {saved} structure artifact(s)")
```
For production monomer runs:
- Use `selected_models: [1, 2, 3, 4, 5]` unless the user requests a smoke test.
- Use `relax_prediction: True` in Python payloads when relaxation is desired;
JSON examples may show `true`.
- State the sequence length caveat: hosted API docs list 1-1000 residues, while
local support-matrix docs list up to 2048 residues on supported hardware.
- If the task is a complex rather than a monomer, redirect to OpenFold3 or
Boltz2.
Treat tiny toy sequences and single-sequence MSAs as API smoke tests, not
quality evidence. For scientific interpretation and validation, read
`references/science.md` and `references/validation.md`.
## Troubleshooting
- `401`: missing, expired, or unauthorized NGC API key.
- `422`: invalid amino-acid characters, sequence too long, malformed A3M, bad
`selected_models`, or malformed mmCIF template object.
- Local `404`: remove `/v1/` from the prediction URL.
- Weak structures: use MSA Search to generate deeper A3M alignments and add
biologically relevant mmCIF templates when appropriate.
- Local startup stalls: first run downloads parameters into `LOCAL_NIM_CACHE`.
Referenced files: 5
openfold3-nim7.08 KB
---
name: openfold3-nim
description: >
Use this skill for OpenFold3, NVIDIA's BioNeMo NIM microservice for biomolecular structure prediction. Invoke whenever the user mentions OpenFold3 or needs protein, protein-ligand, protein-DNA/RNA, or multi-chain complex prediction with the hosted NVIDIA API or local Docker NIM. Covers endpoint choice, auth, request payloads, output artifacts, confidence scores, and local container setup.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# OpenFold3 NIM
Predict biomolecular structures with OpenFold3. It supports proteins, DNA, RNA,
small-molecule ligands, and multi-entity assemblies. Use this `SKILL.md` for
basic hosted/local NIM use; load supplemental files only when the task needs
deeper context:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: purpose, strengths, limitations, and model handoffs.
- `references/parameters.md`: molecule fields, MSAs, templates, samples, tuning.
- `references/validation.md`: artifact checks and scientific sanity checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted URL: `https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict`
- Local URL: `http://localhost:8000/biology/openfold/openfold3/predict`
- Local readiness: `http://localhost:8000/v1/health/ready`
Mode difference: the local prediction path has no `/v1/` prefix. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Auth And Environment
Do not print API keys. Confirm they exist with shell tests, not echoes.
Hosted needs `NGC_API_KEY` in the request header. Local startup needs
`NGC_API_KEY`, or `NVIDIA_API_KEY` as a fallback, plus `LOCAL_NIM_CACHE`.
A repo-root `.env` file may be sourced as a local override before validation.
## Local Docker
Use the official OpenFold3 NIM image and mount `LOCAL_NIM_CACHE` at
`/opt/nim/.cache`. First startup downloads model artifacts and can take several
minutes.
When writing local setup commands, copy the preflight below exactly. Do not
replace it with a simple `: "${NGC_API_KEY:?Set NGC_API_KEY}"` check, do not
drop `NVIDIA_API_KEY`, and do not invent a default `LOCAL_NIM_CACHE`; those
lines are the repo's local NIM env contract. The default single-GPU launch
should show the literal `--gpus "device=0"`; choose a different device only
when the user asks.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm --name openfold3 \
--runtime=nvidia \
--gpus "device=0" \
--shm-size=16g \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/openfold/openfold3:latest
```
Readiness check:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## Request Pattern
Use `requests.post(..., json=payload, timeout=300)`. For local Docker tasks,
set `hosted = False` after the readiness check passes.
```python
import os
import requests
hosted = True
url = (
"https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict"
if hosted
else "http://localhost:8000/biology/openfold/openfold3/predict"
)
headers = {"Content-Type": "application/json"}
if hosted:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
seq = "MKTVRQERLKSIVR"
payload = {
"inputs": [{
"input_id": "prediction_1",
"output_format": "pdb",
"molecules": [{
"type": "protein",
"id": "A",
"sequence": seq,
"diffusion_samples": 1,
"msa": {
"main": {
"a3m": {
"alignment": f">query\n{seq}",
"format": "a3m"
}
}
}
}]
}]
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
Payload gotchas:
- Top level is `{"inputs": [...]}` and OpenFold3 accepts exactly one input.
- `molecules` can contain 1-32 objects with `type`: `protein`, `dna`, `rna`,
or `ligand`.
- Protein/RNA MSAs are optional but, when supplied, `alignment` must start with
a FASTA header such as `>query\nSEQUENCE`.
- Ligands use either `smiles` or `ccd_codes`, for example
`{"type": "ligand", "id": "L", "ccd_codes": "ATP"}`.
- DNA/RNA entities use `sequence`, for example
`{"type": "dna", "id": "B", "sequence": "ATCGATCG"}`.
- `diffusion_samples` is 1-5. `output_format` is `pdb` or `cif`.
## Save And Interpret Output
Save every returned structure as a scientific artifact. Main response path:
`result["outputs"][0]["structures_with_scores"]`.
```python
output = result["outputs"][0]
for i, sample in enumerate(output["structures_with_scores"], start=1):
fmt = sample["format"]
with open(f"openfold3_structure_{i}.{fmt}", "w", encoding="utf-8") as fh:
fh.write(sample["structure"])
print("confidence_score", sample.get("confidence_score"))
print("complex_plddt_score", sample.get("complex_plddt_score"))
print("ptm_score", sample.get("ptm_score"))
print("iptm_score", sample.get("iptm_score"))
print("complex_pde_score", sample.get("complex_pde_score"))
```
Higher `confidence_score`, `complex_plddt_score`, `ptm_score`, and `iptm_score`
are generally better; lower `complex_pde_score` is generally better. Treat toy
or very short sequences as API smoke tests, not meaningful structural biology.
For why and when OpenFold3 is scientifically appropriate, read
`references/science.md`.
## Common Limits
- Inputs per request: 1.
- Molecules per input: 1-32.
- Diffusion samples: 1-5.
- TensorRT path supports shorter sequences; PyTorch path can support longer
sequences, but long inputs need much more GPU memory.
- Sequences over roughly 1800 residues require at least 80 GB GPU memory.
- Local NIM is single-GPU only; choose the target device in the Docker flag.
## Troubleshooting
- `401`: missing, expired, or unauthorized NGC API key.
- `422`: invalid molecule type, invalid sequence characters, bad MSA shape, or
`diffusion_samples` outside 1-5.
- MSA errors: ensure the alignment starts with `>query\n`.
- Local `404`: remove `/v1/` from the prediction URL.
- Local startup stalls: first run may be downloading 10-15 GB of model weights
into `LOCAL_NIM_CACHE`.
- Memory errors: shorten the sequence, reduce samples, or use a larger GPU.
Referenced files: 5
parabricks8.19 KB
---
name: parabricks
description: >-
Route NVIDIA Parabricks pbrun tools, assess GPU/runtime readiness, and provide
version-aware command guidance for FASTQ/BAM processing, RNA-seq, variant
calling, BAM QC, and GVCF workflows. Do NOT use for inspecting or accelerating
whole pipelines — use genomics-workflow-acceleration.
license: CC-BY-4.0 AND Apache-2.0
metadata:
version: "1.1.0"
tags:
- parabricks
- genomics
- nvidia
---
# Parabricks
## Purpose
Use this skill to discover the right NVIDIA Parabricks `pbrun` command, assess
runtime readiness, and generate version-aware command guidance for individual
tools and pipelines.
Do **not** use this skill for whole-workflow inspection, acceleration planning,
or wiring optional GPU branches. For pipeline-level work, use
`genomics-workflow-acceleration`.
## When to Use This Skill
- Which `pbrun` tool fits the user's data and goal
- GPU, driver, Docker, container, storage, or installation readiness
- Command shape, flags, and validation for a specific Parabricks tool
- Troubleshooting a single Parabricks command or tool family
## Prerequisites
Ask for input data type, sequencing technology, reference build, sample
structure, desired output, target Parabricks version/container tag, and runtime
target before recommending commands.
If the user is unsure which tool applies, read
[tool-index.md](references/tool-index.md) first, then load the matching
`references/pbrun-<tool>.md` file.
## Limitations
This skill routes and guides Parabricks commands. It does not install
Parabricks, infer missing sample metadata, guarantee output parity, provide
clinical interpretation, or promise exact runtime without benchmark data.
## Workflow
1. Confirm the Parabricks version or container tag. Verify the current NVIDIA
docs when the user asks for the latest tool list or version-sensitive flags.
2. Classify the request:
- **Runtime** → [runtime-environment.md](references/runtime-environment.md)
- **Tool discovery** → [tool-index.md](references/tool-index.md)
- **Specific command** → matching `references/pbrun-<tool>.md`
3. Collect missing biological and filesystem context before generating commands.
4. Generate conservative Docker commands with explicit mounts, workdir, and
placeholders. Validate paths, indexes, and outputs after command generation.
## Tool Reference Index
Load only the reference file for the selected tool.
| Tool | Reference | Use when |
|------|-----------|----------|
| `applybqsr` | [pbrun-applybqsr.md](references/pbrun-applybqsr.md) | Apply BQSR table to aligned BAM |
| `bam2fq` | [pbrun-bam2fq.md](references/pbrun-bam2fq.md) | BAM → FASTQ conversion |
| `bamsort` | [pbrun-bamsort.md](references/pbrun-bamsort.md) | Standalone BAM sort |
| `bqsr` | [pbrun-bqsr.md](references/pbrun-bqsr.md) | Generate BQSR recalibration table |
| `fq2bam` | [pbrun-fq2bam.md](references/pbrun-fq2bam.md) | Short-read DNA paired FASTQ → BAM/CRAM |
| `fq2bam_meth` | [pbrun-fq2bam_meth.md](references/pbrun-fq2bam_meth.md) | Bisulfite/methylation FASTQ → BAM/CRAM |
| `giraffe` | [pbrun-giraffe.md](references/pbrun-giraffe.md) | Pangenome graph alignment |
| `markdup` | [pbrun-markdup.md](references/pbrun-markdup.md) | Standalone duplicate marking |
| `minimap2` | [pbrun-minimap2.md](references/pbrun-minimap2.md) | Long-read FASTQ alignment |
| `rna_fq2bam` | [pbrun-rna_fq2bam.md](references/pbrun-rna_fq2bam.md) | RNA-seq FASTQ(s) → splice-aware BAM (STAR alignment) |
| `starfusion` | [pbrun-starfusion.md](references/pbrun-starfusion.md) | Fusion detection from chimeric junction input + STAR-Fusion genome library |
| `germline` | [pbrun-germline.md](references/pbrun-germline.md) | GATK-style germline pipeline from FASTQ |
| `deepvariant_germline` | [pbrun-deepvariant_germline.md](references/pbrun-deepvariant_germline.md) | DeepVariant germline pipeline from FASTQ |
| `haplotypecaller` | [pbrun-haplotypecaller.md](references/pbrun-haplotypecaller.md) | Standalone HaplotypeCaller from BAM/CRAM |
| `deepvariant` | [pbrun-deepvariant.md](references/pbrun-deepvariant.md) | Standalone DeepVariant from BAM/CRAM |
| `somatic` | [pbrun-somatic.md](references/pbrun-somatic.md) | Tumor-normal somatic pipeline |
| `mutectcaller` | [pbrun-mutectcaller.md](references/pbrun-mutectcaller.md) | Mutect2-compatible somatic calling |
| `deepsomatic` | [pbrun-deepsomatic.md](references/pbrun-deepsomatic.md) | DeepSomatic-based somatic calling |
| `pacbio_germline` | [pbrun-pacbio_germline.md](references/pbrun-pacbio_germline.md) | PacBio long-read germline |
| `ont_germline` | [pbrun-ont_germline.md](references/pbrun-ont_germline.md) | Oxford Nanopore long-read germline |
| `pangenome_germline` | [pbrun-pangenome_germline.md](references/pbrun-pangenome_germline.md) | Pangenome-aware germline |
| `pangenome_aware_deepvariant` | [pbrun-pangenome_aware_deepvariant.md](references/pbrun-pangenome_aware_deepvariant.md) | Pangenome-aware DeepVariant |
| `prepon` | [pbrun-prepon.md](references/pbrun-prepon.md) | Pangenome-aware preprocessing |
| `postpon` | [pbrun-postpon.md](references/pbrun-postpon.md) | Pangenome-aware post-processing |
| `bammetrics` | [pbrun-bammetrics.md](references/pbrun-bammetrics.md) | Whole-genome coverage/depth metrics |
| `collectmultiplemetrics` | [pbrun-collectmultiplemetrics.md](references/pbrun-collectmultiplemetrics.md) | Multiple Picard/GATK-style alignment metrics |
| `genotypegvcf` | [pbrun-genotypegvcf.md](references/pbrun-genotypegvcf.md) | Joint-genotype GVCF input(s) into VCF |
| `indexgvcf` | [pbrun-indexgvcf.md](references/pbrun-indexgvcf.md) | Index GVCF input |
| `dbsnp` | [pbrun-dbsnp.md](references/pbrun-dbsnp.md) | dbSNP annotation on variant files |
For routing heuristics when multiple tools could apply, see
[tool-index.md](references/tool-index.md).
## Runtime Readiness
For GPU, driver, Docker, container, storage, or installation questions, read
[runtime-environment.md](references/runtime-environment.md) and prefer:
```bash
python3 skills/parabricks/scripts/check_parabricks_runtime.py
```
Add `--path <dir>` for known input/output/tmp paths. Run container probes only
with user consent.
## Command Shape
```bash
docker run --rm --gpus all \
--volume /host/input:/workdir \
--volume /host/output:/outputdir \
--workdir /workdir \
nvcr.io/nvidia/clara/clara-parabricks:<version> \
pbrun <selected-tool> \
<tool-specific-options>
```
Check the version-specific tool reference before finalizing flags.
## Troubleshooting
| Error | Cause | Solution |
|-------|-------|----------|
| Multiple plausible tools | Data type or goal underspecified | Ask for assay, inputs, caller preference, desired output; use tool-index |
| Exact flag requested | Options are version-sensitive | Check the selected tool reference and NVIDIA docs |
| Runtime question | GPU, Docker, drivers, or storage | Use runtime-environment reference and diagnostic script |
| Wrong tool family | Assay or input type unclear | Confirm DNA/RNA/methylation/long-read/pangenome before routing |
| CUDA or memory failure | Runtime not ready or GPU memory constrained | Assess runtime before tuning command flags |
## Guardrails
- Treat command availability and options as version-sensitive.
- Do not infer exact flags from command names alone.
- Do not collapse standalone tools and full pipelines when explaining tradeoffs.
- Do not substitute DNA `fq2bam` for RNA, or germline for somatic callers.
- Do not invent sample names, read groups, reference builds, known-sites files,
model files, graph resources, container tags, or output paths.
- Do not install, upgrade, or modify packages. Label setup commands as user-run.
- Do not claim CPU execution of Parabricks tools.
- Do not claim biological or VCF parity without a comparison run.
- Prefer official NVIDIA docs for exact command syntax and option defaults.
## Key References
- Parabricks tool index:
<https://docs.nvidia.com/clara/parabricks/latest/toolreference.html>
- Output accuracy and compatible CPU software versions:
<https://docs.nvidia.com/clara/parabricks/latest/documentation/tooldocs/outputaccuracyandcompatiblecpusoftwareversions.html>
- Getting started:
<https://docs.nvidia.com/clara/parabricks/latest/gettingstarted.html>
- Overview:
<https://docs.nvidia.com/clara/parabricks/latest/overview.html>
Referenced files: 33
protein-binder-design5.66 KB
--- name: protein-binder-design description: > Orchestrate an end-to-end de novo protein binder design campaign against a protein target by composing BioNeMo NIM skills. Use for binder design, minibinder design, de novo binders, RFdiffusion + ProteinMPNN + Boltz2/OpenFold3 pipelines, epitope/hotspot-targeted design, in-silico binder validation, and ranking designs by interface confidence. license: Apache-2.0 compatibility: "numpy>=1.24; requests>=2.28" allowed-tools: Bash, Read, Write, AskUserQuestion --- # Protein Binder Design (workflow) Run a de novo binder design campaign by composing atomic NIM skills. This skill owns orchestration, handoff contracts, filtering, validation, and the run manifest. It does NOT duplicate per-NIM API details — defer those to each atomic skill's `SKILL.md`. ## Composed skills | Step | Skill | Owns | |---|---|---| | Backbones | `rfdiffusion-nim` | binder backbone PDBs (contigs + hotspots) | | Sequences | `proteinmpnn-nim` | sequences for each backbone | | Co-fold / score | `boltz2-nim` or `openfold3-nim` | binder–target complex + confidence / ipTM | | MSA (optional) | `msa-search-nim` | target A3M for higher-quality folding | The atomic NIM skills are recommended companions (one per NIM, from the BioNeMo NIM skill set). They are **not required**: `references/pipeline.md` carries the concrete request shape for every NIM call, so an agent with NIM access can follow this skill standalone. For endpoints/auth see **Configuration** below. ## Pipeline 1. **Target prep** — get the target PDB + epitope; map epitope/hotspot author residue numbers to RFdiffusion `hotspot_res` strings; optionally build a target MSA with `msa-search-nim`. 2. **Backbones** (`rfdiffusion-nim`) — binder contig + `hotspot_res`; N backbones. 3. **Sequences** (`proteinmpnn-nim`) — k sequences per backbone; drop the native/WT row from `mfasta`. 4. **Co-fold + score** (`boltz2-nim` / `openfold3-nim`) — co-fold binder+target; collect interface confidence (ipTM) and binder pLDDT. 5. **Self-consistency** — CA-RMSD between the RFdiffusion backbone and the predicted binder (`scripts/metrics.py`). 6. **Filter + rank** — apply thresholds; rank survivors; write manifest + CSV. Full handoff contracts, branching, and the cost funnel: `references/pipeline.md`. ## Handoff contracts (the fragile glue) - RFdiffusion `output_pdb` → ProteinMPNN `input_pdb` (inline PDB text). - ProteinMPNN `mfasta` → Boltz2 binder polymer `sequence` (exclude the native/WT row; pair scores only with designed rows). - Epitope author residue numbers → 1-based sequence indices: remap with `scripts/pdb_utils.py:remap_to_seq_index`. RFdiffusion `hotspot_res` uses chain+author strings like `"A50"`; Boltz2 pocket/contacts use 1-based indices. - Boltz2 complex `.cif` → binder chain → self-consistency RMSD vs the backbone. ## Run manifest (reproducibility backbone) Every campaign writes `manifest.json` (+ `candidates.csv`) under a run dir via `scripts/manifest.py`. It records lineage, params, scores, artifacts, filter status, and controls — enabling ranking, resumability, validation, and the final report. Schema and usage: `references/manifest.md`. ## Filters (defaults) - ipTM ≥ 0.8, binder pLDDT ≥ 80, self-consistency RMSD ≤ 2.0 Å. - Override per campaign and record overrides in the manifest `filters`. ## Validation Always run controls and report a **success rate**, not just top scores. Negative controls via `scripts/controls.py` (scrambled sequences); positive controls = published binders re-scored through the same pipeline. Benchmark targets live in `assets/targets.json` (`scripts/registry.py`). Methodology and metric definitions: `references/validation.md`. ## Human-in-the-loop + cost - Confirm target, epitope/hotspots, binder length range, and hosted-vs-local with the user before generating backbones (AskUserQuestion). - Co-folding is the expensive stage: co-fold a capped shortlist, review, then expand. State hosted vs local once and reuse it across all NIM calls. ## Responsible use De novo binder design is dual-use. Decline requests aimed at enhancing pathogen fitness, toxin potency, or bioweapon function; keep designs to legitimate research and therapeutic intent. ## Configuration (NIM access) Each composed NIM is reached over HTTP; choose **hosted** or **local** once and reuse it for every call: - **Hosted** (managed): base URL `https://health.api.nvidia.com/v1/...` per NIM at [build.nvidia.com](https://build.nvidia.com); set `NVIDIA_API_KEY` (sent as `Authorization: Bearer`). Read keys from the env — never hardcode them. - **Local** (self-hosted NGC containers): point each NIM at its local URL (e.g. `http://localhost:8000/...`); local NIMs need no auth header. To **launch** the NIMs yourself (docker run per NIM, persistent caches, health checks, and the GPU **profile‑selection gotcha** — some NIMs (e.g. Boltz2) need `NIM_MODEL_PROFILE` pinned on GPUs that have no bundled profile, while others (RFdiffusion/ProteinMPNN) auto‑select by compute capability): see **`references/local-nim-setup.md`**. Per-NIM paths, request/response schemas, and worked `curl`/Python examples live in `references/pipeline.md`. ## Scripts - `scripts/manifest.py` — campaign manifest (create / load / score / filter / rank / CSV). - `scripts/pdb_utils.py` — PDB parse, chain extract, sequence, residue remap, CA coords. - `scripts/metrics.py` — Kabsch CA-RMSD for self-consistency. - `scripts/controls.py` — scrambled negative controls. - `scripts/registry.py` + `assets/targets.json` — **example** benchmark target registry (illustrative epitopes — verify against the cited structure before a real campaign). Replace with your own targets.
Referenced files: 12
proteinmpnn-nim4.78 KB
---
name: proteinmpnn-nim
description: >
Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Use for ProteinMPNN, inverse folding, sequence design, backbone redesign, fixed chains/residues, omit_AAs, sampling temperature, soluble model, hosted NVIDIA API, local Docker, PDB input, and multi-FASTA output.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# ProteinMPNN NIM
Design protein sequences for a supplied backbone PDB. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: inverse-folding uses, limits, and validation.
- `references/parameters.md`: design controls, fixed positions, sampling.
- `references/validation.md`: FASTA, score, and structure checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`
- Local: `http://localhost:8000/biology/ipd/proteinmpnn/predict`
Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker
For local setup answers, copy the preflight below exactly before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not answer with only
a localhost Python request. This NIM's cache mount is unique:
`/home/nvs/.cache/nim`, not `/opt/nim/.cache`.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run -it \
--runtime=nvidia \
--gpus "device=${NIM_TEST_GPU}" \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/home/nvs/.cache/nim" \
-p 8000:8000 \
nvcr.io/nim/ipd/proteinmpnn:latest
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## Request Pattern
Read PDB content inline; do not send only a file path.
```python
import os
from pathlib import Path
import requests
HOSTED = True
pdb_content = Path("1R42.pdb").read_text()
url = (
"https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
if HOSTED else "http://localhost:8000/biology/ipd/proteinmpnn/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"input_pdb": pdb_content,
"num_seq_per_target": 10,
"sampling_temp": [0.1],
"use_soluble_model": False,
"ca_only": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
Common controls:
- Redesign only chain A: `"input_pdb_chains": ["A"]`.
- Exclude amino acids: `"omit_AAs": ["C"]` or `"omit_AAs": ["M"]`.
- Diversity: `"sampling_temp": [0.1, 0.3, 0.5]` (always a list).
- Solubility bias: `"use_soluble_model": True`.
- Candidate count: `num_seq_per_target` is 1-100.
## Save And Report Output
```python
mfasta = result["mfasta"]
Path("designed_sequences.fa").write_text(mfasta)
# Scores correspond to designed sequences. The mfasta may include a native/WT
# row; do not pair that row with generated-sequence scores.
headers = [line for line in mfasta.splitlines() if line.startswith(">")]
designed_headers = [
h for h in headers if "native" not in h.lower() and "wt" not in h.lower()
]
for header, score in zip(designed_headers, result.get("scores", [])):
print(f"{header} score: {score:.4f}")
```
Validate promising designs by predicting structures with Boltz2 or OpenFold3
and comparing them to the target backbone. For FASTA/score sanity checks, read
`references/validation.md`.
## Limits And Troubleshooting
- Minimum GPU VRAM: about 3 GB.
- `sampling_temp` must be a list, even for one value.
- Empty `mfasta`: check non-empty `input_pdb` and `num_seq_per_target >= 1`.
- PDB parse errors: use valid PDB ATOM records.
- Local URL 404 usually means an accidental `/v1/` prefix.
- Cache mount error: use `/home/nvs/.cache/nim` inside the container.
Referenced files: 5
rfdiffusion-nim5.3 KB
---
name: rfdiffusion-nim
description: >
Run RFDiffusion protein backbone design via NVIDIA NIM. Use for de novo protein backbones, motif scaffolding, binder design, hotspot residues, contigs syntax, diffusion steps, hosted NVIDIA API calls, local Docker deployment, and PDB backbone outputs for ProteinMPNN sequence design.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---
# RFDiffusion NIM
Design protein backbone PDBs for de novo proteins, motif scaffolds, and binders.
Use this `SKILL.md` for first-pass hosted/local usage; load supplemental files
only when needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: design modes, strengths, limits, and handoffs.
- `references/parameters.md`: contigs, hotspots, steps, and seeds.
- `references/validation.md`: PDB, contig, and artifact sanity checks.
- `references/examples.md`: compact hosted/local request patterns.
## Choose Mode
Ask only when context is unclear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/rfdiffusion/generate`
- Local: `http://localhost:8000/biology/ipd/rfdiffusion/generate`
Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
## Local Docker
For local setup answers, copy the preflight below exactly before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not replace it with a
simple `: "${NGC_API_KEY:?Set NGC_API_KEY}"` check, do not invent a cache
default, and do not drop the `NVIDIA_API_KEY` fallback. Default setup is single
GPU `device=0`.
```bash
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"
docker run -it \
--runtime=nvidia \
--gpus "device=0" \
-e NGC_API_KEY \
-v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/ipd/rfdiffusion:2
```
Readiness:
```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```
## Contigs DSL
`contigs` defines what to keep and what to generate.
- `"100"`: generate exactly 100 residues.
- `"80-120"`: generate 80-120 residues.
- `"A25-35"`: keep chain A residues 25-35 from `input_pdb`.
- `"A25-35/0 50-80"`: keep A25-35, insert chain break `/0`, generate 50-80.
Design modes:
- De novo: `contigs="80-120"`; live hosted validation requires a non-empty
`input_pdb` or `input_pdb_asset`, so inline requests should include the dummy
PDB below.
- Motif scaffolding: read `target.pdb`, pass `input_pdb`, use a contig like
`"A25-35/0 50-80"`.
- Binder design: pass target `input_pdb`, contig with target and binder segment,
and `hotspot_res=["A50", "A51", ...]` in ChainResidue string format.
```python
DUMMY_PDB = (
"CRYST1 1.000 1.000 1.000 90.00 90.00 90.00 P 1 1\n"
"ATOM 1 CA ALA A 1 0.000 0.000 0.000 1.00 0.00 C\n"
"END\n"
)
```
## Request Pattern
```python
import os
from pathlib import Path
import requests
HOSTED = True
url = (
"https://health.api.nvidia.com/v1/biology/ipd/rfdiffusion/generate"
if HOSTED else "http://localhost:8000/biology/ipd/rfdiffusion/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
payload = {
"input_pdb": DUMMY_PDB,
"contigs": "80-120",
"diffusion_steps": 50,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
Path("designed_backbone.pdb").write_text(result["output_pdb"])
```
Motif scaffold:
```python
payload = {
"input_pdb": Path("target.pdb").read_text(),
"contigs": "A25-35/0 50-80",
"diffusion_steps": 50,
}
```
Binder design:
```python
payload = {
"input_pdb": Path("target.pdb").read_text(),
"contigs": "A1-100/0 50-100",
"hotspot_res": ["A50", "A51", "A52", "A53", "A54"],
"diffusion_steps": 50,
}
```
## Save And Interpret Output
Save `result["output_pdb"]` as a PDB artifact and report `elapsed_ms` when
present. Generated backbones are not final proteins; feed them to ProteinMPNN
for sequence design, then validate sequences/structures with Boltz2 or
OpenFold3. For PDB and contig checks, read `references/validation.md`.
## Limits And Troubleshooting
- `diffusion_steps`: 1-50; 50 is maximum quality, fewer is faster.
- Single GPU; minimum GPU VRAM is about 12 GB.
- `hotspot_res` uses strings like `"A50"`, not tuples.
- `422` usually means chain IDs in `contigs`/`hotspot_res` do not match
`input_pdb`, a malformed contig, or omitted `input_pdb` for hosted de novo.
- Local URL 404 usually means an accidental `/v1/` prefix.
Referenced files: 5
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 12:00 UTC
- Collection status
- Collected
plugins_6a76572d8f8081918362aa7ff90947fb
Download listing JSON