← Plugin catalog
Scientific Research

NVIDIA BioNeMo Agent Toolkit

NVIDIA BioNeMo v0.1.0

A decade of NVIDIA life-sciences libraries, NIMs, and open models as portable agent skills: structure prediction, docking, generative chemistry, genomics, and de novo protein design. The plugin helps an agent pick the right tool, prepare inputs, run it, and interpret outputs.

Language: English · Automatically detected from descriptions.

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
Apache-2.0
Package author
NVIDIA BioNeMo
Keywords
bionemo, life-sciences, drug-discovery, protein-folding, molecular-docking, generative-chemistry, genomics, protein-design, cuda-x, nim

Declared capabilities

  • Interactive
  • Write

Package observed Sep 30, 2026.

Files & skills

File archives

Plugin package178 files · 427 KBBrowse files →
cuequivariance1 files · 4.9 KBBrowse files →
nvmolkit-usage1 files · 6.36 KBBrowse files →
Skill instructions
boltz2-nim5.25 KB

View saved version →

---
name: boltz2-nim
description: >
  Use Boltz2 NIM for biomolecular structure prediction and binding affinity. Invoke for Boltz2, protein structures, protein-ligand/DNA/RNA complexes, SMILES or CCD ligands, pIC50/IC50 affinity scoring, mmCIF output, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Boltz2 NIM

Predict biomolecular structures and optional ligand affinity. Use this
`SKILL.md` for first-pass hosted/local usage; load supplemental files only when
needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: purpose, strengths, limitations, and handoffs.
- `references/parameters.md`: prediction, sampling, MSA, template, affinity tuning.
- `references/validation.md`: mmCIF, confidence, affinity, and chemistry checks.
- `references/examples.md`: compact hosted/local payload patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/mit/boltz2/predict`
- Local: `http://localhost:8000/biology/mit/boltz2/predict`

Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker

For local setup answers, copy the preflight below before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not invent a cache
default or drop the `.env` load or `NVIDIA_API_KEY` fallback.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm --name boltz2 --gpus all \
  --shm-size=16G \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/mit/boltz2:1.6.0
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

First startup downloads about 30 GB of model weights.

## Request Pattern

```python
import os
import requests

HOSTED = True
url = (
    "https://health.api.nvidia.com/v1/biology/mit/boltz2/predict"
    if HOSTED else "http://localhost:8000/biology/mit/boltz2/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "polymers": [{
        "id": "A",
        "molecule_type": "protein",
        "sequence": "MTEYKLVVVGACGVGKSALTIQLIQNHFVDEYDPT",
    }],
    "recycling_steps": 3,
    "sampling_steps": 50,
    "diffusion_samples": 1,
    "step_scale": 1.638,
    "output_format": "mmcif",
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

Payload essentials:

- Protein polymer: `{"molecule_type": "protein", "sequence": "..."}`.
- DNA/RNA polymer: add another polymer with `molecule_type` `"dna"` or `"rna"`.
- Ligand by SMILES: `{"id": "L1", "smiles": "CC(=O)OC1=CC=CC=C1C(=O)O"}`.
- Ligand by CCD: `{"id": "L1", "ccd": "ATP"}`.
- Affinity: set `"predict_affinity": True` on exactly one ligand; report
  `affinity_pic50`, `affinity_pred_value`, and `affinity_probability_binary`.
- Precomputed A3M MSA goes under the protein polymer. The A3M record uses
  `alignment`, `format`, and `rank`; do not use a stale `data` field.

```python
protein_with_msa = {
    "id": "A",
    "molecule_type": "protein",
    "sequence": "MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT",
    "msa": {"msa_search": {"a3m": {
        "alignment": ">query\nMTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT",
        "format": "a3m",
        "rank": 0,
    }}},
}
```

## Save And Report Output

```python
for i, structure in enumerate(result["structures"], start=1):
    with open(f"structure_{i}.cif", "w", encoding="utf-8") as handle:
        handle.write(structure["structure"])
for i, score in enumerate(result.get("confidence_scores", []), start=1):
    print(f"structure {i} confidence {score:.4f}")
if "affinities" in result:
    for ligand_id, aff in result["affinities"].items():
        print(ligand_id, aff["affinity_pic50"][0], aff["affinity_pred_value"][0], aff["affinity_probability_binary"][0])
```

Save every `.cif` artifact. Visualize in PyMOL, ChimeraX, or UCSF Chimera. For
confidence/affinity sanity checks, read `references/validation.md`.

## Limits And Troubleshooting

- Polymers/request: 12. Ligands/request: 20. Chain length: 4096 residues.
- Affinity prediction supports one ligand per request and adds runtime.
- `422`: invalid sequence, invalid CCD/SMILES, malformed MSA, or multiple
  affinity ligands.
- Local URL/auth: local path has no hosted auth header; wait on `/v1/health/ready`.
- Local startup: use `--gpus all`, `--shm-size=16G`, and the `/opt/nim/.cache` mount.

Referenced files: 5

complexa-binder-design12.5 KB

View saved version →

---
name: complexa-binder-design
description: >
  Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided test-time search (best-of-n, beam search, FK steering, MCTS), select with the internal AF2 reward gate, then INDEPENDENTLY validate each binder by refolding the complex with Boltz2 (default) or OpenFold3 and rank on interface confidence, pLDDT, ipSAE, apo/holo stability, and hotspot contact. Use whenever the user wants de novo binders against a named target, sequence, or PDB, hotspot/epitope-targeted design, Proteina-Complexa / Complexa, or ranked validated binders from one request. Sibling of protein-binder-design (RFdiffusion + ProteinMPNN); this skill uses Proteina-Complexa.
license: Apache-2.0
compatibility: "python>=3.10; numpy>=1.24; gemmi (target prep + Boltz2 templates); pyyaml (target registration)"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Complexa Binder Design (workflow)

From one request — "design binders for `<target>`" — to ranked, **independently
validated** binders. Each returned binder is a **co-designed sequence + predicted
binder–target complex**, gated by interface confidence, by whether the binder
actually contacts the target hotspots, and by **apo/holo stability**.

Generation uses **Proteina-Complexa** (co-designs binder sequence + full-atom
structure together — no inverse-folding step — with reward-guided test-time search).
Validation uses a **different** model family (Boltz2 / OpenFold3), so the headline
confidence is an independent check, not the generator grading its own homework.

> **Upstream model + code (you provide these):**
> - Project page: <https://research.nvidia.com/labs/genair/proteina-complexa/>
> - Code: <https://github.com/NVIDIA-Digital-Bio/Proteina-Complexa> (the `complexa` CLI)
> - Weights (NGC): `nvidia/clara/proteina_complexa`
> - Paper: Didi et al., *Scaling Atomistic Protein Binder Design…*, ICLR 2026.

> **First time on a host? → `references/setup.md`** — full standalone setup with **no
> NIM**: install Proteina-Complexa + download weights, Python deps (`numpy gemmi
> pyyaml`), AF2 **configure-vs-bypass**, optional analyze tools (`foldseek`/`sc`/`dssp`),
> the Boltz2/OF3 validation endpoint, and every env var. Then run
> `bash scripts/check_setup.sh` for a one-shot readiness checklist.

```
Stage 1: Resolve target + hotspots            → target.pdb + hotspots.json   (no GPU)
        ┌──────────────────────────────────────────────────────────────────────┐
        │ repeat until ≥ N validated passers (or a stop cap):                    │
Stage 2 │   Generate (complexa design) → complex .pdb + AF2-reward-gated designs │
Stage 3 │   Validate (Boltz2 default; OF3 optional) → holo+apo + ipTM/ipSAE/     │
        │     pLDDT + apo↔holo RMSD + hotspot contact → passers                  │
        └──────────────────────────────────────────────────────────────────────┘
Stage 4: Report                               → REPORT_<target>_<run>.md (GO/NO-GO)
```

## Composed pieces (read on demand — do not inline)

| Step | Tool | Owns |
|---|---|---|
| Target + hotspots | vendored `science-skills` (UniProt, AFDB) + `scripts/` | structure resolution, evidence-based hotspots, ≤500 crop, preflight |
| Generation | **Proteina-Complexa** `complexa` CLI | co-designed binder seq+structure, AF2-reward gate → `references/complexa-cli.md` |
| Validate / score | `boltz2-nim` (default) or `openfold3-nim` | independent holo+apo refold, ipTM / pLDDT / PAE |
| MSA (target) | `msa-search-nim` or `scripts/fetch_target_msa_colabfold.py` | target A3M for higher-confidence refolds |

If you run inside the Proteina-Complexa repo, its bundled `.claude/skills`
(`complexa-setup`, `complexa-target`, `complexa-design`) can drive the generation
half; this skill adds the automated Stage 1, the independent validation, GO/NO-GO,
and the manifest.

## Stage 1 — resolve target and hotspots (no GPU)

The user gives a target as a **name**, **sequence**, and/or **structure file**.
Resolve exactly **one design-ready structure**, in priority order: (1) experimental
**PDB** (RCSB), (2) **AFDB** model (UniProt → `vendor/science-skills/.../fetch_structure.py`),
(3) **user-provided** file, (4) **fold de novo** (MSA-Search + OpenFold3/Boltz2).
`scripts/pipeline.py:resolve_target_spec`/`resolve_target` automate (1)–(2) from
free text.

**Hotspots** = the target residues the binder should contact — a compact,
surface-exposed, binder-accessible epitope. Resolve in evidence order
(`scripts/hotspot_strategy.py`, `scripts/pdb_interface.py`):

1. **PDB co-complex interface** (gold standard) — interface residues from a structure
   where the target contacts a protein partner.
2. **UniProt functional residues** — `Mutagenesis` + accessible `Active/Binding/Site`,
   **filtered to the extracellular/accessible range** (catalytic/cytoplasmic pockets
   are the wrong surface for a binder and are dropped).
3. **Literature (Paperclip)** — full-text mining when 1–2 are empty
   (`prompts/hotspot_paperclip.md`); structure-confirmed to auto-correct numbering.
4. **Unconditioned** (`[]`) only as a documented last resort.

Then enforce, deterministically:
- **Structure alignment** (`align_hotspots_to_structure`) — drop residues absent from
  the coordinate file; read back the real 3-letter identity (catches UniProt↔PDB
  numbering offsets — never assume equal indices or chain `A`).
- **Epitope sanity** (`_prune_hotspots`) — one compact patch: drop outliers > 30 Å
  from the cluster centroid, cap at 15 residues, prefer ≥ 2.
- **Size budget ≤ 500 residues** (`_crop_target_to_epitope`) — Complexa builds an
  O(n²) pair-feature map over the whole complex, so crop large targets to an epitope
  window (original numbering preserved).

**Preflight (no GPU):** `python3 scripts/preflight_design.py <name|accession> …`
reports the conditioned length, re-aligned hotspots + source, compactness, the ≤500
budget, and a READY / NEEDS-ATTENTION verdict. Review before spending GPU.

## Stage 2 — generate (Proteina-Complexa, open CLI)

Register the target (hotspots + binder length are target-dict-driven), then **use
`complexa generate` (NOT the full `complexa design`)** for the lean, fast path:

```bash
python scripts/complexa_design.py run --task-name <name> --run-name <run> \
    --algorithm best-of-n --num-samples <N> --seed 0 --out <run-dir>
```

`complexa_design.py run` defaults to the **`generate`** verb. With **`best-of-n` + the
AF2 reward** (AF2 params configured via `setup_af2_params.sh` + `AF2_DIR`), the search
**AF2-selects the best candidates during generation** and writes co-designed
**sequence + structure** PDBs to `inference/` — **use the sequence directly, do not
MPNN-redesign it**. Search algorithms: `best-of-n` (default) · `beam-search` ·
`fk-steering` · `mcts`. Overrides + outputs: `references/complexa-cli.md`.

> **Do NOT run the full `complexa design` for this workflow.** Its `evaluate` stage
> **re-folds every design with AF2/RF3/ESMFold (redundant** — best-of-n already
> AF2-selected during search**)** and its `analyze` stage needs `foldseek`/`sc`
> (usually not installed). It is much slower and adds a failure mode. The lean
> `generate` → independent **Boltz2** validation (Stage 3) is the intended path.
>
> No AF2 params? use `--af2-bypass` (`single-pass` + drop the AF2 reward); selection
> then falls entirely to the independent Boltz2 gate (Stage 3). Low-complexity
> (poly-X) sequences are dropped before spending Boltz2.

## Stage 3 — validate (independent refold) + gate

**Validate a capped shortlist, not the whole pool.** Best-of-n produces many
candidates; fold only ~**2× the requested N** (the top ones by the generation/AF2
reward) — validating the entire pool wastes GPU/time and (on hosted Boltz2) trips rate
limits. Point at a local Boltz2 NIM via `$BOLTZ2_URL` (`--endpoint local`) when available.

Per binder run **two** predictions with one refolder (Boltz2 default): **holo**
(binder + target; target MSA, binder single-sequence, `write_full_pae`) and **apo**
(binder alone). One command does it: **`scripts/boltz2_refold.py`** makes the holo
calls (with retry/backoff for rate limits) and chains **`scripts/validate_binders.py`**,
which runs apo + computes the metrics + applies the gate + ranks. Per-chain
conditioning + metric definitions: `references/validation.md`.

**Gate (defaults — every gate must hold):** ipTM ≥ 0.65, complex pLDDT ≥ 0.70, binder
pLDDT ≥ 0.70, apo binder pLDDT ≥ 0.70, **ipSAE_min ≥ 0.45**, apo↔holo binder RMSD
≤ 2.5 Å, ≥ 20% of conditioned hotspots contacted (CB–CB < 13 Å). Record **every**
design (pass *and* fail) with a `failure_reason`. Rank protein binders by interface
confidence (ipTM/ipSAE) + pLDDT + stability — **not** Boltz2 `affinity_pic50`
(ligand-only).

## Bounded-budget loop + report

The deliverable is **the top-N binders ranked by interface confidence** (default 10).
Aim for N that pass the full gate, but **bound the cost**: run **at most 2 generation
rounds**, then **deliver the top-N by score (ipTM, then ipSAE_min) even if fewer than N
clear the strict gate** — keep each design's `pass`/`failure_reason` flag so quality is
still visible. Do **not** keep generating just to chase N strict passes (that is the
single biggest time sink). Stop on N-passed / 2 rounds / budget / a zero-passer round.
One run dir per campaign; `manifest.json` records target, Complexa run config + seeds,
per-design lineage/scores/artifacts, gate status. The report states GO/NO-GO,
requested-vs-achieved N (passed and delivered), ranked binders, and which stop
condition fired. Layout, loop, and report sections: `references/pipeline.md`.

## Configuration

- `COMPLEXA_REPO` — path to your local Proteina-Complexa checkout (the `complexa`
  CLI runs there). Checkpoints via the pipeline YAML or `++ckpt_path=…`. Reward
  weights (`AF2_DIR`, `RF3_CKPT_PATH`/`RF3_EXEC_PATH`) via the repo's `.env`.
- Boltz2 / OpenFold3 endpoints + auth: hosted (`https://health.api.nvidia.com/v1/…`
  + `NVIDIA_API_KEY`) or local (`http://localhost:8000/…`, no auth). The validator
  takes `--endpoint hosted|local`; the key is read from `NVIDIA_API_KEY`/`NGC_API_KEY`
  (or `--env-file`). Never hardcode hosts/keys.
- `COMPLEXA_OUTPUTS` — run-output root (default `./outputs`).

## Scripts & assets

- `references/setup.md` + `scripts/check_setup.sh` — standalone (no-NIM) install guide
  and a one-shot environment readiness check.
- `scripts/pipeline.py` — orchestrator (Stage-1 resolution + open-CLI generation +
  AF2 gate + scoring); `score_existing` and `full` modes.
- `scripts/preflight_design.py` — no-GPU target/hotspot/size planner.
- `scripts/hotspot_strategy.py`, `scripts/pdb_interface.py` — evidence-based hotspots.
- `scripts/complexa_design.py` — thin `complexa design` driver + output extraction.
- `scripts/setup_af2_params.sh` — download AF2-Multimer params (public, no auth) +
  create the `params/` layout, for reward-guided search (`best-of-n`, etc.).
- `scripts/boltz2_refold.py` — Stage-3 **holo** Boltz2 refolds (retry/backoff + throttle)
  → `validation/raw/`, then chains `validate_binders.py`.
- `scripts/validate_binders.py` — apo Boltz2 + scoring, ipSAE, apo↔holo RMSD,
  hotspot contact, gating, ranking → `ranked_binders.json`. Needs the Dunbrack **ipSAE**
  script: `bash scripts/fetch_ipsae.sh` (MIT; fetched, not bundled — see `vendor/ipsae/`).
- `scripts/fetch_target_msa_colabfold.py`, `scripts/pdb_to_boltz_template_cif.py` —
  target MSA / structural-template helpers for validation.
- `vendor/science-skills/` — DeepMind UniProt + AFDB tooling (Apache-2.0) for Stage 1.
- `prompts/hotspot_paperclip.md` — literature-mining hotspot fallback.
- `assets/targets.json` — example registered targets (use `complexa target add` for your own).

## Responsible use

De novo binder design is dual-use. Decline requests aimed at enhancing pathogen
fitness, toxin potency, or bioweapon function; keep designs to legitimate research
and therapeutic intent.

## See also

- `protein-binder-design` — same goal via RFdiffusion + ProteinMPNN (BioNeMo NIMs).
- Proteina-Complexa docs: `README.md`, `docs/INFERENCE.md`, `docs/CONFIGURATION_GUIDE.md`,
  `docs/EVALUATION_METRICS.md`, and its bundled `.claude/skills/` in the repo above.

Referenced files: 32

complexa-design16.2 KB

View saved version →

---
name: complexa-design
description: >
  End-to-end Proteina-Complexa design pipeline driver. Reach for this skill whenever
  the user wants to "design a binder", "design binders for X", "run complexa
  design", "de novo binder", "PDL1 binder", "TrkA binder", "design proteins for
  target", "protein binder design", "ligand binder", "design a small-molecule
  binder", "ATP-binding protein", "AME motif scaffolding", "scaffold a motif
  near a ligand", "motif + ligand design", "enzyme scaffolding", "flow matching
  protein design", "beam-search binder", "FK steering", "MCTS protein design",
  "refold with AF2", "refold with RF3", or wants success rates, interface pAE,
  scRMSD, or
  FoldSeek diversity from a single command. This is the scientific anchor of
  the skill set: it drives `complexa design <pipeline>` from target picking to
  manifest emission and tells the user how many designs passed.
compatibility: "complexa CLI installed (pip install -e .); .env populated; 1x CUDA GPU >=40GB VRAM (A100/H100/L40S); 24 CPUs; ~50GB disk"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Complexa Design Skill

Drive the full four-stage `complexa design` pipeline: generate (flow matching +
search) -> filter (top-N by reward) -> evaluate (refold with AF2/RF3) ->
analyze (success rate, FoldSeek/MMseqs diversity). Pick the right pipeline
config for the design intent, validate the run upfront so the user does not
discover a missing ckpt mid-folding, run it, and emit a replayable manifest +
per-design success CSV.

## What this skill enables

- Protein binder design for protein targets (AF2 reward + ColabDesign refold).
- Ligand binder design for small-molecule targets (RF3 reward + RF3 refold).
- AME motif scaffolding with ligand context (motif + ligand features, RF3).
- Search-based optimization: single-pass, best-of-n, beam-search, fk-steering, mcts.
- Refold backends: ColabDesign (AF2), RF3, Boltz2, ESMFold (fast iteration).
- Pass-rate + diversity analysis with per-`result_type` thresholds.

## Step 1: Pre-flight

Always run the shared preflight before launching a design — generation needs the
GPU and the right checkpoint, evaluation needs AF2/RF3 weights and tool
binaries. Bail early if the host cannot run the chosen pipeline.

```bash
bash .claude/skills/_shared/scripts/preflight.sh
```

Read `./complexa_setup/preflight.json` and bail if any of these are missing for
the chosen pipeline:

- `gpu.available: false` -> all pipelines fail.
- `gpu.vram_gb < 40` -> generation OOMs at default `batch_size: 16`; lower to 8.
- `ckpts.complexa[.ckpt]` -> required for protein binder.
- `ckpts.complexa_ligand[.ckpt]` -> required for ligand binder.
- `ckpts.complexa_ame[.ckpt]` -> required for AME.
- `env.AF2_DIR` missing -> protein binder default eval (`colabdesign`) fails.
- `env.RF3_CKPT_PATH` or `env.RF3_EXEC_PATH` missing -> ligand binder / AME default eval (`rf3_latest`) fails.

If a ckpt is missing, point at `complexa-setup` and have the user run
`complexa download --complexa-<variant>` first.

## Step 2: Pick the pipeline

Complexa has **one default pipeline (protein binder) and two extensions**
(ligand binder, AME / enzyme). The pipeline is selected entirely
by the `configs/search_*_pipeline.yaml` you pass to `complexa design` — each
YAML pins its own model checkpoint, autoencoder, targets dict, default reward,
and default refold backend. Switching pipelines is just "swap the config path
and the target name comes from a different dict".

### Default — protein binder

```bash
complexa design configs/search_binder_local_pipeline.yaml \
    ++run_name=pdl1_v1 ++generation.task_name=02_PDL1
```

Use this when the user says "design a binder for X", "PDL1 binder",
"de novo binder", "design proteins for a target", etc. — i.e. the target is
a protein surface. The config pins `complexa.ckpt` + `complexa_ae.ckpt`, reads
targets from `configs/targets/targets_dict.yaml`, rewards with AF2
(`af2folding`), inverse-folds with SolubleMPNN, and evaluates with
ColabDesign / AF2. **If the user did not specify, this is what they want.**

### Extension A — ligand binder (small-molecule pocket)

```bash
complexa design configs/search_ligand_binder_local_pipeline.yaml \
    ++run_name=v11_v1 ++generation.task_name=39_7V11_LIGAND \
    ++metric.binder_folding_method=rf3_latest
```

Switch to this when the user says "ligand binder", "small-molecule pocket",
"SMILES target", "ATP-binding protein", or names a target ending in
`_LIGAND` / from the FAD / SAM / OQO / 7V11 / etc. families. The config
**also activates LoRA** (`r=32`, `lora_alpha=64`) which is required for the
released ligand checkpoint — leave the `lora:` block alone.

### Extension B — AME / motif + ligand (enzyme scaffolding)

```bash
complexa design configs/search_ame_local_pipeline.yaml \
    ++run_name=ame_chm ++generation.task_name=M0096_1chm
```

Switch to this when the user says "scaffold a motif near a ligand",
"active-site design", "enzyme scaffolding", "AME", or names a target like
`M0024_1nzy`, `M0096_1chm` (the `M####_<pdb>` AME task naming). The config sets
`env_vars.USE_V2_COMPLEXA_ARCH=True` (the CLI runner injects it into the
subprocess) and uses both `MotifFeatures` and `LigandFeatures`. Default search
is `single-pass`; switch to `best-of-n` only if you also enable the
`CompositeRewardModel` (commented out in `ame_generate.yaml`).

### Pipeline cheat sheet — what changes when you switch

| Knob | Protein binder (default) | Ligand binder | AME (enzyme) |
|---|---|---|---|
| **Pipeline YAML** | `configs/search_binder_local_pipeline.yaml` | `configs/search_ligand_binder_local_pipeline.yaml` | `configs/search_ame_local_pipeline.yaml` |
| **Model ckpt** | `complexa.ckpt` | `complexa_ligand.ckpt` | `complexa_ame.ckpt` |
| **Autoencoder ckpt** | `complexa_ae.ckpt` | `complexa_ligand_ae.ckpt` | `complexa_ame_ae.ckpt` |
| **Targets dict** | `configs/targets/targets_dict.yaml` | `configs/targets/ligand_targets_dict.yaml` | `configs/design_tasks/ame_dict_v2.yaml` |
| **Task-name pattern** | `<NN>_<NAME>` (e.g. `02_PDL1`, `22_DerF21`) | `<NN>_<PDB>_LIGAND` (e.g. `39_7V11_LIGAND`) | `M####_<pdb>` (e.g. `M0096_1chm`) |
| **`USE_V2_COMPLEXA_ARCH`** | (unset → v1) | (unset → v1) | `"True"` (set in YAML) |
| **LoRA** | (none) | required (`r=32, alpha=64`) | required |
| **Default search algo** | `best-of-n` | `best-of-n` | `single-pass` |
| **Reward model** | AF2 (`af2folding`) | RF3 (`rf3folding`) | `null` (no reward at default) |
| **Inverse folder** | `soluble_mpnn` | `ligand_mpnn` | `ligand_mpnn` |
| **Default refold backend** | `colabdesign` (AF2) | `rf3_latest` | `rf3_latest` |
| **Analysis `result_type`** | `protein_binder` | `ligand_binder` | `motif_ligand_binder` |
| **Required ckpts (`complexa download`)** | `--complexa --all` (AF2 in community) | `--complexa-ligand --all` (RF3 in community) | `--complexa-ame --all` (RF3 in community) |

For the full per-pipeline breakdown (reward weights, success thresholds,
analysis modes), see [reference/pipelines.md](reference/pipelines.md).

## Step 3: Gather parameters

Use AskUserQuestion to fill in the four parameters that vary every run. Default
to sensible production settings if the user has no preference.

- **Target name** — must be a key in the relevant dict (`targets_dict.yaml` for
  protein binder, `ligand_targets_dict.yaml` for ligand, `ame_dict_v2.yaml` for
  AME). If the user names a target that is not in the dict, hand off to
  `complexa-target` to add it first.
- **Run name** — a short identifier appended to the output dir (e.g. `pdl1_v1`).
- **Search algorithm** — default to `beam-search` with `beam_width=8` and
  `n_branch=4` for production. Use `single-pass` for a quick smoke test.
- **Evaluation refold backend** — protein binder defaults to `colabdesign`
  (AF2); ligand/AME default to `rf3_latest`. Use `esmfold` for fast iteration
  (worse but seconds per sample).

## Step 4: Validate

Validate before running. This is cheap (seconds) and catches missing ckpts,
missing env vars, unknown override keys, and missing target entries — all of
which would otherwise abort the pipeline mid-evaluation after hours of
generation.

```bash
complexa validate design configs/search_binder_local_pipeline.yaml \
    ++generation.task_name=02_PDL1 \
    ++metric.binder_folding_method=colabdesign
```

The validator returns non-zero on failure and prints a pass/fail report.
Re-run the command with the suggested overrides until it returns clean.

## Step 5: Run the pipeline

`complexa design` is the right tool for the full 4-stage run — it orchestrates
`generate → filter → evaluate → analyze` as sequential subprocesses with a
shared run name, log dir, and multi-GPU split (see `run_design_pipeline` in
`src/proteinfoundation/cli/cli_runner.py`). Re-implementing that manually
loses the per-stage log routing and progress prints.

Use `++` (forced) Hydra overrides; they apply to all stages. The minimal
production protein-binder invocation:

```bash
complexa design configs/search_binder_local_pipeline.yaml \
    ++run_name=pdl1_v1 \
    ++generation.task_name=02_PDL1 \
    ++generation.search.algorithm=beam-search \
    ++generation.search.beam_search.beam_width=8 \
    ++metric.binder_folding_method=colabdesign
```

For ligand binder / AME, swap the pipeline YAML and target name per Step 2's
cheat sheet — every other override above is pipeline-agnostic and can be
reused as-is.

Add `--verbose` to stream logs to the terminal instead of `./logs/`. The skill
does not poll progress — the user re-invokes if they want a status; point them
at `complexa status` and `./logs/design_pipeline_*/`.

### Direct module invocation (debug fallback)

To debug a single stage without pipeline orchestration, invoke the underlying
Hydra module directly. `complexa generate CONFIG` is just a logged subprocess
wrapper around this:

```bash
python -m proteinfoundation.generate \
    --config-path "$(realpath configs)" \
    --config-name search_binder_local_pipeline \
    ++run_name=debug_pdl1 \
    ++generation.task_name=02_PDL1
```

Same pattern for `proteinfoundation.{filter,evaluate,analyze}`. Use this when
you want to attach `ipdb`, run under `nsys`, or skip the pipeline log dir.
Prefer `complexa generate/filter/evaluate/analyze` for normal one-shot runs
(you get logging and parallel job splitting for free).

**AME-specific gotcha (when running AME with RF3 refold)**: RF3 will try to
complete missing atoms on the ligand based on its CCD code, which produces
shape errors in RMSD calculations and the wrong structure. Before RF3 sees the
PDB, rename the ligand residue to `L:0` so RF3 treats it as a generic ligand
and skips atom completion. If your AME generation outputs already encode the
ligand this way (the canonical Complexa pipeline does), no extra step is
needed; otherwise patch each PDB with `atomworks.io`:

```python
from atomworks.io import load_any, to_pdb_file
atom_array = load_any("my_design.pdb")[0]
ligand_mask = atom_array.chain_id == "A"
atom_array.res_name[ligand_mask] = "L:0"
to_pdb_file(atom_array, "my_design_rf3_ready.pdb")
```

Same applies if you're piping AME outputs into the `complexa-evaluate-pdbs`
skill with an RF3 backend. Skip the rename for AF2/colabdesign refold or for
non-AME pipelines.

Wall-clock at default (`nsteps=400`, `beam_width=8`, `batch_size=16`, 100
designs, colabdesign eval) is ~30–120 minutes on a single A100/H100.

## Step 6: Collect results

Outputs land in two directories. Surface both:

```bash
ls ./inference/${CONFIG_STEM}_${TASK}_*${RUN_NAME}/   # generated PDBs + filter
ls ./evaluation_results/${RUN_NAME}/                  # per-design CSV + analysis
```

Read the combined results CSV and summarize:

```bash
ls ./evaluation_results/*/binder_results_*_combined.csv
ls ./evaluation_results/*/motif_binder_results_*_combined.csv  # AME
ls ./evaluation_results/*/res_filter_*_pass_*.csv              # success rate
ls ./evaluation_results/*/res_div_foldseek_*.csv               # FoldSeek diversity
```

Pull the success rate from `res_filter_binder_pass_*.csv`, the per-design
metrics (interface pAE, pLDDT, scRMSD) from the combined CSV, and FoldSeek
TM-score diversity from `res_div_foldseek_*.csv`. Report top-N designs by
i_pAE (protein binder) or min_ipAE (ligand binder).

## Step 7: Emit manifest

Drop a JSON manifest beside the results so the run is replayable. The shared
helper captures the command, config, git SHA, and pointers to the result CSVs.

```bash
python3 .claude/skills/_shared/scripts/write_manifest.py \
    --output-dir ./evaluation_results/${RUN_NAME} \
    --command "complexa design configs/search_binder_local_pipeline.yaml ++run_name=${RUN_NAME} ++generation.task_name=${TASK}" \
    --skill complexa-design \
    --out ./run_manifest.json
```

Surface the manifest path and the result CSV to the user.

## Most-common overrides

The 10 overrides that cover ~90% of runs. Full reference (every key, type,
default) is in [reference/overrides.md](reference/overrides.md).

| Override | Default | What it controls |
|----------|---------|------------------|
| `++generation.task_name=<name>` | (per config) | Which target / AME task to design for |
| `++run_name=<str>` | (config stem) | Output dir suffix and CSV tag |
| `++generation.search.algorithm=beam-search` | `best-of-n` (binder/ligand), `single-pass` (AME) | Search strategy |
| `++generation.search.beam_search.beam_width=8` | `4` | Beam-search width (more = better designs, slower) |
| `++generation.args.nsteps=200` | `400` | Diffusion steps (fewer = faster, lower quality) |
| `++generation.dataloader.batch_size=8` | `16` (binder/ligand/AME) | Drop to 8 on a 40GB GPU |
| `++generation.filter.filter_samples_limit=500` | `1000` | Top-N samples to keep after filtering |
| `++metric.binder_folding_method=esmfold` | `colabdesign` (binder), `rf3_latest` (ligand/AME) | Evaluation refold backend |
| `++metric.num_redesign_seqs=8` | `2` | ProteinMPNN/LigandMPNN/SolubleMPNN sequences per design |
| `++aggregation.success_thresholds.i_pAE.threshold=10.0` | `7.0` (protein binder) | Loosen / tighten success criteria |

## Hardware requirements

| Resource | Minimum | Recommended |
|----------|---------|-------------|
| GPU | 1x CUDA GPU, 40 GB VRAM | A100 / H100 / L40S, 80 GB VRAM |
| CPUs | 16 | 24 (the `ncpus_` default in every pipeline config) |
| Disk | 50 GB at `./inference/` + `./evaluation_results/` | 200 GB for sweep runs |
| RAM | 32 GB | 64 GB+ |

Typical wall-clock for 100 designs, `beam_width=8`, default `nsteps=400`:

- Protein binder + colabdesign refold: ~60–120 min on 1x A100/H100.
- Ligand binder + RF3 refold: ~90–180 min (RF3 dominates).
- AME + RF3 refold: ~120–240 min.
- Any pipeline + ESMFold refold: ~30–60 min (fast iteration).

Bumping `gen_njobs=2` and `eval_njobs=2` halves wall-clock on a 2-GPU host. See
`.claude/skills/_shared/reference/hardware.md` for per-pipeline VRAM tables.

## Troubleshooting (common cases)

| Symptom | Cause | Fix |
|---------|-------|-----|
| `CUDA out of memory` in generate | `batch_size: 16` too big on 40GB GPU | `++generation.dataloader.batch_size=8` |
| `CUDA out of memory` in evaluate | AF2 / RF3 batched too aggressively | `++eval_njobs=1` and `++metric.num_redesign_seqs=2` |
| `InterpolationKeyError: AF2_DIR` | colabdesign eval but `.env` does not set `AF2_DIR` | Set `AF2_DIR` in `.env` or `++metric.binder_folding_method=esmfold` |
| `InterpolationKeyError: RF3_CKPT_PATH` | RF3 eval but RF3 not installed | `complexa download --all` or switch eval backend |
| `KeyError: 'task_name' not in target_dict_cfg` | Target not in `targets_dict.yaml` / `ligand_targets_dict.yaml` / `ame_dict_v2.yaml` | Use `complexa-target` skill to add it |
| 0 designs pass success thresholds | Defaults too strict for this target | Loosen via `++aggregation.success_thresholds.*` |

For the full list (chain-ID mismatches, hotspot residues, ligand residue
renaming for RF3, missing inverse-folding models, etc.) see
[reference/troubleshooting.md](reference/troubleshooting.md).

---

For per-pipeline details (which model, reward, inverse folding, evaluation
backend, `result_type`, LoRA settings, `USE_V2_COMPLEXA_ARCH` toggle), see
[reference/pipelines.md](reference/pipelines.md).

For the full override reference (every `generation.*`, `metric.*`,
`aggregation.*` key with type, default, example, and effect), see
[reference/overrides.md](reference/overrides.md).

For all troubleshooting cases (cause + fix + source), see
[reference/troubleshooting.md](reference/troubleshooting.md).

Referenced files: 3

complexa-evaluate-pdbs13.2 KB

View saved version →

---
name: complexa-evaluate-pdbs
description: >
  Standalone evaluation of an existing PDB directory with Proteina-Complexa.
  Use this skill whenever the user wants to "evaluate PDB files", "re-fold these
  designs", "compute interface pAE", "compute i_pLDDT for a folder",
  "run AF2 / RF3 / ESMFold on my designs", "score binder candidates",
  "designability of this folder", "scRMSD for designs", "motif RMSD for these
  PDBs", "complexa analysis", "complexa evaluate from a PDB directory",
  "evaluate from pdb dir", or score third-party outputs (BindCraft, AlphaProteo,
  RFdiffusion, hand-curated decoys). It picks the correct `evaluate_*.yaml`
  config, wires `++dataset.pdb_dir` and the folding backend, runs
  `complexa analysis` (the evaluate → analyze chain), parses the result CSV,
  reports pass-rates against the right `result_type` thresholds, and emits a
  replayable `eval_manifest.json`. Reach for this skill before hand-rolling
  refolding scripts.
compatibility: "complexa CLI installed (pip install -e .); CUDA GPU; AF2_DIR (colabdesign) or RF3_CKPT_PATH+RF3_EXEC_PATH (rf3_latest); ESMFold weights for monomer paths"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Complexa Evaluate-PDBs Skill

Score a directory of pre-existing PDB files against the same metrics Proteina-Complexa uses internally. Wraps `complexa analysis <evaluate_config> ++sample_storage_path=<dir>`: the CLI runs the `evaluate` step (refold + interface metrics + monomer metrics) and then the `analyze` step (success thresholds, diversity, pass-rate CSVs). Do **not** run `complexa generate` here — the inputs already exist.

## What this skill enables

- Re-fold a directory of designed PDBs with AF2 (`colabdesign`), RF3 (`rf3_latest`), ESMFold (`esmfold`), or Boltz2 (`boltz2_default`).
- Compute binder interface metrics: `i_pAE`, `min_ipAE`, `i_pTM`, `pLDDT`, binder/complex scRMSD.
- Compute monomer **designability** (ProteinMPNN-redesigned scRMSD) and **codesignability** (original sequence refold scRMSD).
- For motif inputs: motif RMSD (CA + all-atom), motif-region designability/codesignability, sequence recovery.
- Aggregate into per-PDB CSVs plus pass-rate summaries using the default thresholds for the `result_type`.

## Step 1: Pre-flight

Always check GPU / disk / tool binaries before launching a refold job. RF3 and ColabDesign-AF2 are large.

```bash
bash .claude/skills/_shared/scripts/preflight.sh
```

Surface from `preflight.json`:

- `gpu.available` and `gpu.vram_gb` — colabdesign/RF3 need ≥40 GB; ESMFold tolerates ≥24 GB.
- `env.missing_required` — must include the keys for the chosen folding backend:
  - `colabdesign` → `AF2_DIR`
  - `rf3_latest` → `RF3_CKPT_PATH`, `RF3_EXEC_PATH`
  - `esmfold` → ESMFold weights resolvable
- `tools.{foldseek,mmseqs}` — required by `aggregation.compute_diversity` / `compute_mmseqs_diversity` (both default `true`).

If any required key is missing, route the user to the `complexa-setup` skill.

## Step 2: Identify the design type

Like `complexa design`, evaluation has one default flow (protein binder) and two extensions (ligand binder, AME). The evaluate config you pass to `complexa analysis` decides everything else (which metrics, which refolder defaults, which thresholds the analyze step applies).

### Default — protein binder

```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml \
    ++sample_storage_path=/abs/path/to/pdbs \
    ++dataset.task_name=02_PDL1 \
    ++result_type=protein_binder \
    ++metric.binder_folding_method=colabdesign \
    ++metric.inverse_folding_model=soluble_mpnn \
    ++run_name=eval_pdl1_af2
```

Use this when the user's PDBs are protein-binder designs (multi-chain, binder is the last chain) or third-party outputs from BindCraft / AlphaProteo / RFdiffusion. Pulls thresholds for `protein_binder` (`i_pAE * 31 ≤ 7.0`, `pLDDT ≥ 0.9`, `scRMSD_ca < 1.5 Å`).

### Extensions — pick the matching config

| Design type | Use the protein-binder default? | Evaluate config | Analyze config | `result_type` | Default backend |
|---|---|---|---|---|---|
| Protein binder | **Yes (default)** | `configs/evaluate_from_pdb_dir.yaml` | `configs/analyze.yaml` | `protein_binder` | `colabdesign` (AF2) |
| Ligand binder (binder + small-molecule) | Same evaluate config, swap 3 overrides | `configs/evaluate_from_pdb_dir.yaml` | `configs/analyze.yaml` | `ligand_binder` | `rf3_latest` |
| AME / motif + ligand (enzyme outputs) | No — needs motif-aware config | `configs/evaluate_ame_from_pdb_dir.yaml` | `configs/analyze_motif_binder.yaml` | `motif_ligand_binder` | `rf3_latest` |
| Motif protein binder (standalone) | No — no `_from_pdb_dir` variant | `configs/evaluate_motif_binder.yaml` + `++input_mode=pdb_dir` | `configs/analyze_motif_binder.yaml` | `motif_protein_binder` | `colabdesign` / `rf3_latest` |

**Extending to ligand binder** (same evaluate config as default, three override swaps):

```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml \
    ++sample_storage_path=/abs/path/to/pdbs \
    ++dataset.task_name=39_7V11_LIGAND \
    ++result_type=ligand_binder \
    ++metric.binder_folding_method=rf3_latest \
    ++metric.inverse_folding_model=ligand_mpnn \
    ++run_name=eval_v11_rf3
```

**Extending to AME** (different config; ligand auto-completion gotcha — see Step 4):

```bash
complexa analysis configs/evaluate_ame_from_pdb_dir.yaml \
    ++sample_storage_path=/abs/path/to/pdbs \
    ++dataset.task_name=M0096_1chm \
    ++run_name=eval_ame_chm
```

See `reference/eval_configs.md` for the full matrix (every `result_type`, every threshold default, every supported folding backend, motif-protein-binder variant).

## Step 3: Gather inputs (AskUserQuestion)

Ask in one batched `AskUserQuestion`:

1. **`pdb_dir`** — absolute path to the directory of PDBs to evaluate.
2. **Design type** — protein binder / ligand binder / AME (motif + ligand).
3. **Folding backend** — `colabdesign` (AF2, protein binders), `rf3_latest` (ligand / AME), `boltz2_default` (alternative).
4. **Target / task name** — must match a key in `configs/targets/targets_dict.yaml`, `configs/targets/ligand_targets_dict.yaml`, or `configs/design_tasks/ame_dict_v2.yaml`. Required for binder + AME evaluation (needed to identify the target reference and, for AME, the motif contigs).
5. **AME-only**: confirm ligand residue name is already renamed to `L:0` in every PDB (see Troubleshooting). If not, do that rename first.

## Step 4: Run evaluate → analyze

Prefer `complexa analysis` (the evaluate→analyze chain) — it reuses the same config for both steps and writes a single log dir.

```bash
# Protein binder PDB dir, AF2 refold
complexa analysis configs/evaluate_from_pdb_dir.yaml \
  ++sample_storage_path=/abs/path/to/pdbs \
  ++dataset.task_name=02_PDL1 \
  ++metric.binder_folding_method=colabdesign \
  ++metric.inverse_folding_model=soluble_mpnn \
  ++result_type=protein_binder \
  ++run_name=eval_pdl1_af2
```

For ligand binders flip `binder_folding_method=rf3_latest`, `inverse_folding_model=ligand_mpnn`, `result_type=ligand_binder`. For AME use `configs/evaluate_ame_from_pdb_dir.yaml` — see `reference/eval_configs.md` for full worked examples.

If you need to inspect output between stages, run them separately. The configs above are shared between `evaluate` and `analyze`:

```bash
complexa evaluate configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=/abs/path/to/pdbs ++run_name=eval_pdl1_af2
complexa analyze  configs/evaluate_from_pdb_dir.yaml ++run_name=eval_pdl1_af2
```

Dry-run first if the user is unsure (no GPU work happens; the planned file walk + invocation prints):

```bash
complexa analysis configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=/abs/path/to/pdbs ++dryrun=true
```

### Direct module invocation (debug fallback)

`complexa evaluate` / `analyze` are subprocess wrappers around the Hydra
modules with logging + parallel job splitting bolted on. To attach a debugger
or run under a profiler, invoke the module directly:

```bash
python -m proteinfoundation.evaluate \
    --config-path "$(realpath configs)" \
    --config-name evaluate_from_pdb_dir \
    ++sample_storage_path=/abs/path/to/pdbs \
    ++dataset.task_name=02_PDL1 \
    ++metric.binder_folding_method=colabdesign \
    ++run_name=eval_debug
```

For normal one-shot runs prefer `complexa analysis` — you get the shared log
dir and a single replayable invocation, instead of having to thread the same
overrides through two `python -m` calls.

## Step 5: Parse results

Output lands under `./evaluation_results/${run_name}/`:

- Per-PDB metrics CSV — `*_results_*.csv` (one row per input PDB × `sequence_types`).
- Pass-rate summaries — written by the `analyze` step (e.g. `res_designability.csv`, `res_filter_ligand_pass_*.csv`, `success_criteria_*.json`).
- Diversity output — FoldSeek/MMseqs2 cluster files when `aggregation.compute_diversity=true` (default).

Summarize to the user:

- Per-PDB row count and number of successful designs vs total.
- Default-threshold pass rate by `result_type` (e.g. for `protein_binder`: `i_pAE*31 <= 7.0 AND pLDDT >= 0.9 AND scRMSD_ca < 1.5`).
- Top 5 designs by primary metric (`i_pAE` for protein, `min_ipAE` for ligand, `motif_rmsd_pred_all` for motif binders).

## Step 6: Emit manifest

Capture the resolved invocation + outputs for replay.

```bash
python3 .claude/skills/_shared/scripts/write_manifest.py \
  --output-dir ./evaluation_results/${run_name} \
  --command "complexa analysis configs/evaluate_from_pdb_dir.yaml ++sample_storage_path=<dir> ++dataset.task_name=<task> ++metric.binder_folding_method=<backend> ++run_name=<run>" \
  --skill complexa-evaluate-pdbs \
  --out ./eval_manifest.json
```

The manifest pins: resolved config, git SHA, ckpt SHA-256s, the result CSV paths, and the user-stated `result_type` for replay.

## Most common overrides

| Override                                       | Effect                                                                |
|------------------------------------------------|-----------------------------------------------------------------------|
| `++sample_storage_path=<dir>`                  | The directory of PDBs to evaluate (required).                         |
| `++dataset.task_name=<name>`                   | Target / AME task name (binders, AME). Resolves target PDB + contigs. |
| `++metric.binder_folding_method=<backend>`     | `colabdesign` / `rf3_latest` / `boltz2_default`.                      |
| `++metric.inverse_folding_model=<model>`       | `protein_mpnn` / `soluble_mpnn` / `ligand_mpnn`.                      |
| `++metric.sequence_types=[self,mpnn,mpnn_fixed]` | Which sequence flavors to refold.                                   |
| `++metric.num_redesign_seqs=N`                 | ProteinMPNN/LigandMPNN redesign count.                                |
| `++metric.compute_pre_refolding_metrics=true`  | Add bioinformatics/TMOL/HBPLUS metrics on the input structures.       |
| `++metric.keep_folding_outputs=true`           | Save the refolded PDBs (large, but useful for inspection).            |
| `++result_type=<type>`                         | Override default thresholds: `protein_binder` / `ligand_binder` / `motif_protein_binder` / `motif_ligand_binder`. |
| `++aggregation.success_thresholds.<…>`         | Tighten or loosen specific thresholds (see `reference/eval_configs.md`). |
| `++eval_njobs=N`                               | Parallel GPUs for the evaluate step.                                  |
| `++dryrun=true`                                | Plan without running any folding.                                     |
| `++file_limit=N`                               | Cap input PDBs (handy for first-pass smoke tests).                    |

## Hardware

- **GPU**: ≥1 CUDA GPU. AF2 (`colabdesign`) and RF3 (`rf3_latest`) need ≥40 GB VRAM (A100/H100/L40S). ESMFold runs on ≥24 GB. Multi-GPU via `++eval_njobs=N`.
- **CPU/disk**: 24 CPUs default (`ncpus_: 24`). Each refolded PDB + intermediate output is ~1–5 MB; `keep_folding_outputs=true` can balloon to tens of GB for thousands of inputs.
- See `_shared/reference/hardware.md` for per-backend wall-clock and VRAM tables.

## Troubleshooting

- **`Error: Config file not found`** — paths are relative to the repo root; `cd` to the repo before invoking `complexa analysis`.
- **`compute_motif_binder_metrics=True` but `result_type=protein_binder`** — `result_type` and the underlying `compute_*_metrics` must agree. Use `configs/evaluate_ame_from_pdb_dir.yaml` for AME inputs rather than mutating `evaluate_from_pdb_dir.yaml`.
- **RF3 shape errors on AME PDBs** — RF3 tries to auto-complete the ligand atoms from CCD. Rename the ligand residue to `L:0` in every input PDB before evaluation; see the snippet in `README.md` (`atom_array.res_name[ligand_mask] = "L:0"`).
- **Diversity step fails (`foldseek not found`)** — `FOLDSEEK_EXEC` not on `PATH`. Either fix `.env` (preferred) or disable: `++aggregation.compute_diversity=false ++aggregation.compute_mmseqs_diversity=false`.
- **All pass-rates are 0%** — check the `binder_folding_method` matches the target type (RF3 for ligand, AF2 for protein) and that `++dataset.task_name` resolves to the correct reference PDB (`complexa target show <name>` to verify).

## Reference

Full evaluate/analyze config matrix, every supported `result_type`, per-threshold defaults, and three worked examples (protein binder / ligand binder / AME): see `reference/eval_configs.md`.

Referenced files: 1

complexa-setup13.7 KB

View saved version →

---
name: complexa-setup
description: >
  First-time setup, environment configuration, and model-weight installation for
  Proteina-Complexa. Reach for this skill whenever the user says "set up complexa",
  "install complexa", "configure my .env", "first-time setup", "what models do I
  have installed", "what's in my .env", "download model weights", "download
  Complexa / AF2 / RF3 / ProteinMPNN / LigandMPNN / ESM2 / ESMFold checkpoints",
  "preflight my GPU", "verify environment", "complexa init", "complexa download",
  "complexa download --status", "complexa validate env", or any time a fresh
  checkout needs to be made runnable. This is the first skill to run on a new
  clone — it drives `complexa init`, `complexa download`, and `complexa validate
  env` end-to-end, edits the required `.env` keys, picks the right runtime (UV
  vs Docker), and emits a replayable setup artifact.
compatibility: "complexa CLI installed (pip install -e .); bash 4+; nvidia-smi optional"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Complexa Setup Skill

Drive the three steps a fresh Proteina-Complexa checkout needs before any
design run: create `.env`, fetch model weights, and sanity-check the env.
Probe the host for GPU / disk / tool binaries first so the user does not
discover a missing dependency mid-pipeline. End with a JSON setup artifact the
user (or a future agent) can re-read instead of re-deriving state.

## CLI vs direct file-edit — pick the cheapest path per step

| Step | Preferred path | Why |
|---|---|---|
| `.env` creation (Step 2) | **File edit** (`cp .env_example .env` + 3 line swaps) or `complexa init` | `complexa init` is a thin wrapper around `cp + 3 regex swaps` (`_swap_runtime_in_env` in `cli_runner.py`). Either path works; pick CLI for new humans, direct edit for agents. |
| `.env` value edits (Step 3) | **File edit** (StrReplace `LOCAL_CODE_PATH=…` etc.) | No CLI for this — the values are user-specific paths. |
| Download model weights (Step 4) | **CLI** (`complexa download --…`) | Dispatches to `env/download_startup.sh` (~1000 lines of bash with NGC URLs, retries, checksum-style skip-if-present). Don't try to replicate. |
| Validate env (Step 5) | **CLI** (`complexa validate env`) or `test -f .env && test -d $DATA_PATH` | CLI prints a nicer report; the manual check is one-liner-safe. |
| Validate full design config (after picking a pipeline) | **CLI** (`complexa validate design CONFIG`) | Non-trivial Hydra defaults traversal + ckpt + env-var checks; not worth replicating. |

## What this skill enables

- A correctly-shaped `.env` for either UV or Docker runtime.
- Model checkpoints (Complexa protein/ligand/AME plus community models) downloaded to known paths.
- A `preflight.json` snapshot of the host (GPU, disk, .env, ckpts, tool binaries).
- A `run_manifest.json` capturing exactly which `complexa init` + `complexa download` invocations were used (replay-friendly).
- A pass/fail report from `complexa validate env` with clear next-step hints.

## Step 1: Pre-flight check

Always run the shared preflight before touching the environment. It does not
require `.env` to exist — it falls back to defaults — and it tells you whether
the host can run Complexa at all.

```bash
bash .claude/skills/_shared/scripts/preflight.sh
```

The script writes `./complexa_setup/preflight.json`. Read it and surface:

- `gpu.available` — if `false`, design / evaluate steps will fail; warn the user.
- `gpu.vram_gb` — Complexa needs ≥40 GB (A100/H100/L40S).
- `disk.free_gb` at `CKPT_PATH` — minimum ~50 GB for the full Complexa + community model set.
- `env.missing_required` — anything listed here must be edited in `.env` before validation passes.
- `tools.{foldseek,mmseqs,dssp,hbplus,sc}.exists` — missing tools degrade evaluation but do not block generation.

## Step 1b: Build the Python environment (only if `.venv/` is missing)

The `complexa` CLI is installed inside the project's Python environment, not on
the system path by default. On a **fresh clone**, the `.venv/` directory does
not yet exist and `complexa init` will fail with `command not found`. Build the
UV venv before anything else:

```bash
test -d .venv || ./env/build_uv_env.sh   # first-time UV build
source .venv/bin/activate
which complexa                            # sanity check: should point inside .venv
```

Skip this step if `which complexa` already resolves — that means a previous
build is still good. The Docker runtime skips it entirely; the venv lives
inside the container image instead. If the user said "I just cloned" or you
see no `.venv/` next to `pyproject.toml`, run the build script — `complexa init`
without a venv produces a confusing `command not found` rather than an obvious
"build the venv first" error.

## Step 2: Create `.env`

Pick the runtime. UV is the default and faster to start; Docker is required on
Ubuntu 20.04 or systems with GLIBC mismatches.

Use AskUserQuestion if it is not obvious from context:

> "Which runtime do you want to configure? `uv` (recommended, faster) or `docker` (use if you do not have a UV venv built locally)?"

### Path A: file edit (preferred for agents)

`complexa init` only does three things — copy `.env_example` → `.env` and
swap the `COMPLEXA_RUNTIME=` line plus the `UV_*` ↔ `DOCKER_*` prefixes on the
tool/data/cache path block. You can do the same with `cp` + StrReplace and skip
the CLI:

```bash
cp .env_example .env
# Then StrReplace these lines in .env:
#   COMPLEXA_RUNTIME=uv          → COMPLEXA_RUNTIME=<runtime>
#   FOLDSEEK_EXEC=${UV_FOLDSEEK_EXEC}  → FOLDSEEK_EXEC=${DOCKER_FOLDSEEK_EXEC}   (and same for RF3_EXEC_PATH, SC_EXEC, HBPLUS_EXEC, MMSEQS_EXEC, DSSP_EXEC, TMOL_PATH)
#   DATA_PATH=${LOCAL_DATA_PATH} → DATA_PATH=${DOCKER_DATA_PATH}                  (and same for CACHE_DIR, CKPT_PATH)
```

Skip the prefix swap entirely if you're staying on UV (the `.env_example`
already targets UV).

### Path B: CLI

```bash
complexa init                    # UV runtime (default)
complexa init --runtime docker   # Docker runtime
complexa init --force            # Recreate .env from .env_example (drops any edits)
```

If `.env` already exists and `--force` is not passed, only the runtime-dependent
lines are swapped — user edits in Step 3 are preserved across runtime flips.

### Verify either way

```bash
test -f .env && echo "OK: .env present" || echo "MISSING"
grep -E '^COMPLEXA_RUNTIME=' .env
```

## Step 3: Edit .env

No CLI for this — Step 2 only set the runtime; you still need to write your
machine-specific paths into `.env` by hand (StrReplace or your editor). The two
absolutely-required edits are:

```bash
LOCAL_CODE_PATH=/absolute/path/to/protein-foundation-models
LOCAL_DATA_PATH=/absolute/path/to/PFM_data
```

Everything else (cache, ckpts, community-model dirs, tool binaries) is derived
from `LOCAL_CODE_PATH` by default and only needs editing if you have a
non-standard layout. For the full table — every key, what it controls, what
fails if it is missing — see [reference/env_keys.md](reference/env_keys.md).

Quick decision table for the four edits most users make:

| Key | Default | Set this if |
|-----|---------|-------------|
| `LOCAL_CODE_PATH` | placeholder | Always — required |
| `LOCAL_DATA_PATH` | `/path/to/PFM_data` | Always — required, points at target PDBs |
| `HF_TOKEN` | placeholder | You need ESMFold or gated HF models |
| `WANDB_API_KEY` | placeholder | You want training runs logged to W&B |

## Step 4: Download checkpoints

**Always use the CLI here.** `complexa download` dispatches to
`env/download_startup.sh` (~1000 lines of bash with NGC URLs, retries, and
skip-if-present logic across ~6 community-model families). Rolling your own
wget loop is a recipe for partial downloads and wrong destination paths.

Ask which models the user actually needs — downloading everything is ~100+ GB.
Pick from the three Complexa variants and the community-model set. Each
Complexa variant unlocks exactly one `complexa design` pipeline; AF2 / RF3
inside the community-model set are what `evaluate` (and reward-guided search)
need at run time.

| Flag | What it downloads | Unlocks pipeline | Destination | Approx size |
|------|-------------------|------------------|-------------|-------------|
| `--complexa` | Complexa protein-binder model + AE (`complexa.ckpt`, `complexa_ae.ckpt`) | **Protein binder** (default) — `configs/search_binder_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-ligand` | Ligand-binder model + AE (`complexa_ligand.ckpt`, `complexa_ligand_ae.ckpt`) | Ligand binder — `configs/search_ligand_binder_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-ame` | AME motif-scaffolding model + AE (`complexa_ame.ckpt`, `complexa_ame_ae.ckpt`) | AME (enzyme) — `configs/search_ame_local_pipeline.yaml` | `./ckpts/` | ~3 GB |
| `--complexa-all` | All three Complexa variants | All three pipelines | `./ckpts/` | ~9 GB |
| `--all` | All community models (ProteinMPNN + LigandMPNN + AF2 + ESM2 + ESMFold + RF3) | Needed by **evaluate / reward**: AF2 (protein binder), RF3 (ligand binder + AME), MPNNs (inverse folding for every pipeline). | `./community_models/` | ~50 GB |
| `--everything` | Complexa + community + optional (Boltz2 / Protenix) | Everything plus alternative refold backends | both | ~100+ GB |
| `--status` | Show install state — does not download | (none) | (none) | n/a |

**Minimum download per pipeline:**

- Protein binder (default): `complexa download --complexa --all`
- Ligand binder: `complexa download --complexa-ligand --all`
- AME / enzyme: `complexa download --complexa-ame --all`
- All three: `complexa download --everything`

For the full per-model destination breakdown and per-flag NGC sources, see
[reference/downloads.md](reference/downloads.md).

Pick the smallest invocation that covers the user's goal, then run:

```bash
complexa download --complexa            # protein binder only
complexa download --complexa-all        # all three Complexa variants
complexa download --all                 # community models only
complexa download --everything          # everything
```

Without arguments, `complexa download` launches an interactive wizard — prefer
explicit flags in agent mode.

Verify what landed:

```bash
complexa download --status
```

The status output groups by family (Complexa / community / optional) and prints
"Installed" or "Missing" per ckpt. Re-run the specific flag if a ckpt is
flagged missing.

## Step 5: Validate

Final check that `.env` is loadable and the required paths resolve:

```bash
complexa validate env
```

`validate env` checks: (1) `.env` exists, (2) `DATA_PATH` is set and points at
an existing directory. It does not check ckpt files — those are checked by
`complexa validate design <config>` once you have a pipeline config picked.

Common failures and fixes:

| Symptom | Cause | Fix |
|---------|-------|-----|
| `.env file: No .env file found` | `complexa init` not run | Run `complexa init` first |
| `DATA_PATH: Not set in .env` | Placeholder not edited | Edit `LOCAL_DATA_PATH` in `.env` |
| `DATA_PATH: Directory not found` | Path edited but does not exist on disk | `mkdir -p $LOCAL_DATA_PATH` or copy target data there |
| Hydra error `InterpolationKeyError: AF2_DIR` | Reward/eval config wants AF2 but `.env` does not define it | Download AF2 weights or remove AF2 from the config |

## Step 6: Emit setup artifact

Drop a JSON manifest in `./complexa_setup/` so the user has a single file
describing the resulting state. The shared helper writes it for you:

```bash
mkdir -p ./complexa_setup
python .claude/skills/_shared/scripts/write_manifest.py \
    --kind setup \
    --runtime "$(grep -E '^COMPLEXA_RUNTIME=' .env | cut -d= -f2)" \
    --preflight ./complexa_setup/preflight.json \
    --out ./complexa_setup/run_manifest.json
```

Surface the resulting files to the user:

```bash
ls -la ./complexa_setup/
```

Expected contents:

```
complexa_setup/
├── preflight.json      # GPU / disk / .env / ckpt / tool snapshot
└── run_manifest.json   # init + download invocations + git SHA + runtime
```

## Hardware requirements

| Resource | Minimum | Recommended |
|----------|---------|-------------|
| GPU | 1× CUDA GPU, ≥24 GB VRAM | A100 / H100 / L40S, 40–80 GB VRAM |
| CUDA | 12.0 | 12.4+ |
| Disk (CKPT_PATH) | 50 GB | 150 GB (covers `--everything`) |
| RAM | 16 GB | 64 GB+ |
| OS | Ubuntu 22.04+ (UV) | Ubuntu 22.04+ or Docker on any host |

Ubuntu 20.04 throws GLIBC errors with the UV runtime — use `complexa init
--runtime docker` on those hosts. See `.claude/skills/_shared/reference/hardware.md`
for per-pipeline (binder vs ligand vs AME) requirements.

## Troubleshooting

| Symptom | Cause | Fix |
|---------|-------|-----|
| `complexa: command not found` | Package not installed in active env | `source .venv/bin/activate` then `pip install -e .` |
| `complexa init` says `.env_example not found` | Running outside repo root | `cd` to the project root (where `.env_example` lives) |
| `.env_example not found. Cannot initialize .env.` | Not in project root | `cd` into `protein-foundation-models/` and retry |
| `complexa download` fails on NGC URL | Behind firewall / no internet | Configure a proxy for the download script, or download the model `.ckpt`s manually from the NGC pages linked in the main `README.md` and drop them into `./ckpts/` |
| `complexa download --status` shows ckpts present but `validate` fails | `.env` `CKPT_PATH` points elsewhere | Either move ckpts or edit `LOCAL_CHECKPOINT_PATH` in `.env` |
| GLIBC error on import | Ubuntu 20.04 with UV runtime | Re-run `complexa init --runtime docker` and use `./env/docker-ops.sh run` |

---

For the full `.env` reference (every key, defaults, failure modes), see
[reference/env_keys.md](reference/env_keys.md).

For the full download flag matrix, NGC URLs, and destination layout, see
[reference/downloads.md](reference/downloads.md).

Referenced files: 2

complexa-sweep12.4 KB

View saved version →

---
name: complexa-sweep
description: Use this skill whenever the user wants to run a parameter sweep over a Proteina-Complexa design pipeline — cartesian-product hyperparameter scans, Pareto search over generation/reward/evaluation knobs, or any "compare configurations" workflow. Trigger phrases include "sweep beam width", "sweep nsteps", "hyperparameter sweep", "parameter scan", "scan beam_width and temperature", "compare configurations", "find the best generation params", "what's the optimal nsteps", "Pareto search for binder quality vs wall-clock", "complexa sweep", "tune Complexa", "ablate the reward weights", "configs/sweeps", "--sweeper", "run beam_width.yaml". This is the only skill that owns sweeper YAML authoring, cartesian-product expansion, and per-config result ranking.
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# complexa-sweep

Run cartesian-product parameter sweeps over Proteina-Complexa design pipelines. Pick or author a sweeper YAML in `configs/sweeps/`, expand it to N inference configs with `script_utils/generate_inference_configs.py`, loop `complexa design` over those configs, then aggregate per-config success metrics into a ranked summary CSV plus a manifest.

> **Important:** The `complexa design` CLI does **NOT** accept `--sweeper` directly. Sweeps are driven by `script_utils/generate_inference_configs.py`, which writes one `configs/inference_configs/inf_{idx}_{run_name}.yaml` per sweep combination. You then loop `complexa design` over those generated configs.

## What this skill enables

- Pick an existing sweeper YAML from `configs/sweeps/` (`beam_width`, `bb_ca_temperature`, `search_replicas`, `example`).
- Author a new sweeper YAML with arbitrary dot-notation axes (cartesian product).
- Generate N inference + evaluation config pairs with `script_utils/generate_inference_configs.py`.
- Loop `complexa design` over the generated configs (one at a time on a single GPU, or in parallel on a multi-GPU host by sharding the config list across `CUDA_VISIBLE_DEVICES`).
- Walk per-config output directories and parse the analyze-step CSV from each.
- Emit `sweep_summary.csv` (one row per config: axis values + success rate + mean iPAE + diversity) and `sweep_manifest.json`.
- Identify the best config by success rate and the Pareto frontier (wall-clock vs success).

## Step 1: Pre-flight

```bash
bash .claude/skills/_shared/scripts/preflight.sh
```

Read `./complexa_setup/preflight.json`. A sweep multiplies GPU time by the number of configs. **Before launching, confirm the cost with the user**:

> "This sweep produces N configs × ~M minutes per config ≈ TOTAL GPU-hours. OK to proceed? (y / reduce / cancel)"

If `gpu.available=false`, stop — sweeps are not feasible on CPU.

## Step 2: Pick the pipeline + target

Use the same dialogue as `complexa-design` — do **not** duplicate it here. See [`.claude/skills/complexa-design/SKILL.md`](../complexa-design/SKILL.md) Step 2 ("Pick the pipeline") and Step 3 ("Gather parameters"). Capture:

- `pipeline_config_name` — e.g. `search_binder_local_pipeline` (default), `search_ligand_binder_local_pipeline`, `search_ame_local_pipeline`.
- `task_name` — e.g. `02_PDL1`, `22_DerF21`, `39_7V11_LIGAND`. Passed as `--override generation.task_name=<task>`.
- `run_name` — short tag for output dir naming.

## Step 3: Pick or author the sweeper YAML

Sweeper YAMLs live in `configs/sweeps/`. Each key is a dot-notation Hydra path; each value is a list. The cartesian product becomes N configs.

### Canned sweepers

| File | Axis | Values | Configs |
|---|---|---|---|
| `configs/sweeps/beam_width.yaml` | `generation.search.beam_search.beam_width` | 1, 2, 4, 8 | 4 |
| `configs/sweeps/bb_ca_temperature.yaml` | `generation.model.bb_ca.simulation_step_params.sc_scale_noise` | 0.1, 0.4 | 2 |
| `configs/sweeps/search_replicas.yaml` | `generation.search.best_of_n.replicas` | 1, 4, 16, 64 | 4 |
| `configs/sweeps/example.yaml` | beam_width × nsteps | (2,4) × (200,400) | 4 |

If one matches the user's intent, use it as-is. Otherwise author a new file.

### Authoring a new sweeper

Minimal multi-axis example (saved to `configs/sweeps/my_sweep.yaml`):

```yaml
# 3 beam widths × 2 nsteps = 6 configs
generation.search.beam_search.beam_width:
  - 2
  - 4
  - 8

generation.args.nsteps:
  - 200
  - 400
```

Rules (from `script_utils/generate_inference_configs.py:load_sweeper_file`):

- Top-level mapping only. Keys are dot-notation Hydra paths.
- Values must be **lists**. A scalar is auto-wrapped into a single-element list (which pins a value without adding a dimension).
- Cartesian product: total configs = product of list lengths. Two 4-value axes = 16 configs; budget accordingly.
- If a key appears in both the sweeper file and an `--override`, the override wins and that axis collapses.

See [reference/sweep_axes.md](reference/sweep_axes.md) for the full catalogue of swept keys (typical ranges, cost multipliers, what improves/regresses).

### Dry-run preview before generating

Always confirm the config count first:

```bash
python script_utils/generate_inference_configs.py \
    --config_name search_binder_local_pipeline \
    --sweeper configs/sweeps/my_sweep.yaml \
    --override generation.task_name=22_DerF21 \
    --run_name my_sweep \
    --dryrun
```

The output lists every axis + value list and prints `DRY RUN — would generate N config pair(s)`.

## Step 4: Generate configs + loop `complexa design`

Once the dry-run looks right, drop `--dryrun` to materialize `inf_{idx}_{run_name}.yaml` + `eval_{idx}_{run_name}.yaml` pairs under `configs/inference_configs/` and `configs/eval_configs/`. Then loop `complexa design` over the inference configs.

```bash
# 1. Generate one inf_*.yaml + one eval_*.yaml per combination.
python script_utils/generate_inference_configs.py \
    --config_name search_binder_local_pipeline \
    --sweeper configs/sweeps/my_sweep.yaml \
    --override generation.task_name=22_DerF21 \
    --run_name my_sweep

# 2. Loop complexa design over each generated inference config.
for cfg in configs/inference_configs/inf_*_my_sweep.yaml; do
    complexa design "$cfg" || echo "FAILED: $cfg"
done
```

This serialises the sweep on a single GPU — be honest with the user that wall-clock = N × per-run.

### Multi-GPU host (optional speed-up)

On a host with K GPUs, shard the config list K ways and launch one loop per GPU in parallel. Each `complexa design` call uses exactly one GPU at default `gen_njobs=1` / `eval_njobs=1`, so pinning via `CUDA_VISIBLE_DEVICES` keeps them from colliding:

```bash
CONFIGS=(configs/inference_configs/inf_*_my_sweep.yaml)
N=${#CONFIGS[@]}; K=4   # 4 GPUs
for gpu in $(seq 0 $((K-1))); do
    (
        for i in $(seq $gpu $K $((N-1))); do
            CUDA_VISIBLE_DEVICES=$gpu complexa design "${CONFIGS[$i]}" \
                || echo "FAILED on GPU $gpu: ${CONFIGS[$i]}"
        done
    ) &
done
wait
```

Drop `gen_njobs` / `eval_njobs` overrides into the loop only if you intentionally want each `complexa design` invocation to consume multiple GPUs (and you have a way to keep them out of each other's way).

## Step 5: Collect results

Each swept config writes to its own `./inference/inf_{idx}_{run_name}/` directory; the analyze step writes `./evaluation_results/eval_{idx}_{run_name}/results_*.csv`. After the sweep finishes:

```bash
ls -d ./inference/inf_*_my_sweep/
ls ./evaluation_results/eval_*_my_sweep/results_*.csv
```

For each config, parse the analyze CSV (one row per generated binder). Standard columns used for ranking: `i_pae`, `i_plddt`, `sc_rmsd`, `binder_seq`, `passes_filter` (bool). If `passes_filter` is missing, derive a success flag with the user's chosen thresholds (defaults: `i_pae < 10`, `i_plddt > 0.7`, `sc_rmsd < 2.0` — confirm with user).

## Step 6: Rank configs

Emit `sweep_summary.csv` to the run directory. One row per config:

| Column | How to compute |
|---|---|
| `config_id` | The `{idx}` from `inf_{idx}_{run_name}` |
| `<axis_1>`, `<axis_2>`, ... | The swept value at this combination (read from the per-config `inf_*.yaml`) |
| `n_samples` | Row count in the analyze CSV |
| `success_rate` | `passes_filter.mean()` |
| `mean_i_pae` | `i_pae.mean()` (lower = better) |
| `mean_i_plddt` | `i_plddt.mean()` (higher = better) |
| `diversity_score` | Unique sequence count / `n_samples` (or use TM-score clustering if available) |
| `wall_clock_min` | From the per-config log timestamps |

Then report:

- **Best config** = argmax of `success_rate`. Tie-break on `mean_i_pae` (lower).
- **Pareto frontier** over (`wall_clock_min`, `success_rate`): a config is on the frontier iff no other config is both faster AND has higher success rate.

Print the best config + the frontier to the terminal. Save the full table to `sweep_summary.csv`.

## Step 7: Emit manifest

Capture the resolved invocation + outputs for replay. The shared helper takes a single `--output-dir` (it walks for CSVs and pulls the Hydra config from the run's `.hydra/`), so call it against the best-config's evaluation directory and pass the parent `complexa design` command for that run:

```bash
BEST_ID=4   # from Step 6 ranking
python3 .claude/skills/_shared/scripts/write_manifest.py \
    --output-dir ./evaluation_results/eval_${BEST_ID}_my_sweep \
    --command "python script_utils/generate_inference_configs.py --config_name search_binder_local_pipeline --sweeper configs/sweeps/my_sweep.yaml --override generation.task_name=22_DerF21 --run_name my_sweep && for cfg in configs/inference_configs/inf_*_my_sweep.yaml; do complexa design \"\$cfg\"; done" \
    --skill complexa-sweep \
    --out ./sweep_runs/my_sweep/sweep_manifest.json
```

Alongside the manifest, save the ranked `sweep_summary.csv` from Step 6 to `./sweep_runs/my_sweep/sweep_summary.csv` and surface both paths to the user.

## Recommended sweep recipes

| Symptom | Sweep this | Why |
|---|---|---|
| Quality not good enough | `beam_width × nsteps` (use `example.yaml` as a starting point) | Both raise compute → quality; find the cheapest combination that lands. |
| Too slow / want speed-up | `nsteps` downward (e.g. `[100, 200, 400]`) with fixed `beam_width=4` | Find the smallest nsteps that retains success rate. |
| Mode collapse / low diversity | `generation.model.bb_ca.simulation_step_params.sc_scale_noise` (use `bb_ca_temperature.yaml`) | Higher noise → more diverse backbones. |
| Reward over-fitting (high reward, bad metrics) | `generation.reward_model.reward_models.af2folding.reward_weights.{i_pae, plddt}` ratio | Re-balance composite reward. |
| Want statistical robustness on one config | `search_replicas.yaml` (`best_of_n.replicas`) | Same config, more samples → tighter success-rate estimate. |
| Algorithm shoot-out | `generation.search.algorithm` over `[single-pass, best-of-n, beam-search, fk-steering]` | Compare search regimes; pin everything else. |

## Hardware

Total GPU-time for a sweep = `N_configs × per_run_GPU_time`. A single `search_binder_local_pipeline` run on one A100 is roughly 30–90 min (binder length + nsteps dependent). A 4-axis × 4-value sweep = 256 configs × ~60 min = ~256 GPU-hours — at that scale, plan on a multi-GPU host with the Step-4 sharding pattern.

Refer to [`.claude/skills/_shared/reference/hardware.md`](../_shared/reference/hardware.md) for the per-run baseline + VRAM minima.

## Troubleshooting

| Symptom | Cause | Fix |
|---|---|---|
| `Sweeper file not found` from `generate_inference_configs.py` | Path resolved from wrong CWD | Use a path relative to the repo root; or pass an absolute path. |
| `Sweeper YAML must be a mapping` | Top-level YAML is a list or scalar | Rewrite as `key: [v1, v2]` mapping. |
| `No configs were generated` | One of the value lists is empty `[]` | Sweeper file has a `key: []` line — add at least one value. |
| `complexa design` rejects `--sweeper` | Confusion between entrypoints | The CLI does not accept `--sweeper`. Use `generate_inference_configs.py` first, then loop `complexa design` over the generated configs. |
| Override silently collapses a sweep axis | `--override key=v` shadowed a sweep key | Drop either the override OR the matching key from the sweeper file. |
| One config in the loop fails, sweep keeps going | The `\|\| echo "FAILED…"` in Step 4 swallows the error | Re-run the failed `inf_*.yaml` standalone; check the per-config log under `./logs/`. Skip failed `config_id` when ranking. |

---

For per-axis reference (typical ranges, cost, what gets better/worse), see [reference/sweep_axes.md](reference/sweep_axes.md).

For the user-facing sweep system overview (config generation, output layout), see [`docs/SWEEP.md`](reference/SWEEP.md).

Referenced files: 2

complexa-target14.6 KB

View saved version →

---
name: complexa-target
description: Use this skill whenever the user wants to add, register, edit, list, show, or validate a Proteina-Complexa design target for any pipeline — protein binder (default), ligand binder, or AME / enzyme scaffolding. Triggers include "add a target", "define a new target for binder design", "register a hotspot", "set up a PDL1 binder target", "ligand binder pocket", "SMILES target", "AME task", "enzyme motif", "M0024_1nzy", "complexa target add", "complexa target show", "configure target X", "what targets are available", "where do hotspots live", "what does target_input mean", "chain-spec syntax", "binder length range", or any question about `configs/targets/{,ligand_}targets_dict.yaml` and `configs/design_tasks/ame_dict_v2.yaml`. Also covers `complexa validate target`. This is the only skill that touches the three targets dict files.
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# complexa-target

Add or edit a design target in Proteina-Complexa. Targets live in **three YAML files**, one per `complexa design` pipeline:

- `configs/targets/targets_dict.yaml` — protein binder (**default pipeline**)
- `configs/targets/ligand_targets_dict.yaml` — ligand binder
- `configs/design_tasks/ame_dict_v2.yaml` — AME / enzyme scaffolding

The `complexa target` CLI only manages the first two. AME tasks use an extended schema (the same core fields as ligand targets plus `contig_atoms` for hand-curated per-residue motif atom selections) and are file-edit-only.

## Preferred path: edit the YAML directly

`complexa target add` is a thin wrapper around "load YAML → build dict → append a block" (see `src/proteinfoundation/cli/target_manager.py:add_target_cli` and `append_target_to_dict`). For agentic use, **just edit the targets dict directly** — the schema is short, the existing entries are great copy templates, and you skip 14 CLI flags / SMILES shell-escaping. Use the CLI only if you explicitly want its YAML auto-quoter or its overwrite-prompt safety.

The skill therefore presents both paths:

| Step | Direct file edit (preferred) | CLI |
|---|---|---|
| Look at existing targets | Read `configs/targets/{,ligand_}targets_dict.yaml` (or `configs/design_tasks/ame_dict_v2.yaml` for AME) | `complexa target list` (protein/ligand only) |
| Look at one target | Read the dict and grep for the key | `complexa target show NAME` |
| Add a new target | Append a YAML block (Step 3a) | `complexa target add ...` (Step 3b) |
| Verify the PDB resolves on disk | — *(no shortcut, this checks Hydra defaults)* | `complexa validate target CONFIG --target NAME` |

`complexa target` only sees the two protein-style dicts (`targets_dict.yaml` and `ligand_targets_dict.yaml`). For AME tasks (which use a different schema), file-edit is the only path — see Step 1 below.

## What this skill enables

- Register a **protein target** (chain + residue range + hotspots) in `configs/targets/targets_dict.yaml`.
- Register a **ligand target** (PDB pocket + 3-letter code + SMILES) in `configs/targets/ligand_targets_dict.yaml`.
- Resolve the right **AME task name** (e.g. `M0024_1nzy`, `M0096_1chm`) for the AME pipeline — these are not added via `complexa target add` (see Step 1).
- Verify a target resolves to a real PDB on disk with `complexa validate target`.
- Emit a replayable artifact (`target_definition.yaml`) for downstream design runs.

## Step 1: Decide the target type

Targets live in **three different files**, one per `complexa design` pipeline. Pick the file by working out which pipeline the user will run, then add the entry to that file. Each file is consumed by exactly one pipeline.

| User intent | Pipeline | Targets dict file | This skill? |
|---|---|---|---|
| Bind a protein surface (PD-L1, IFNAR2, TNF-α, …) — **default** | `configs/search_binder_local_pipeline.yaml` | `configs/targets/targets_dict.yaml` | Yes — Step 2 (protein) |
| Bind a small-molecule pocket (FAD, SAM, OQO, …) | `configs/search_ligand_binder_local_pipeline.yaml` | `configs/targets/ligand_targets_dict.yaml` | Yes — Step 2 (ligand) |
| AME / enzyme scaffolding (motif + ligand, `M####_<pdb>` names) | `configs/search_ame_local_pipeline.yaml` | `configs/design_tasks/ame_dict_v2.yaml` | Partial — see "AME tasks" below |

### AME tasks (enzyme pipeline)

Similar story — AME tasks live in `configs/design_tasks/ame_dict_v2.yaml` under `motif_target_dict_cfg:` with their own schema (`source`, `target_filename`, `ligand`, `contig_atoms`, `binder_length`, `use_bonds_from_file`). The `contig_atoms` string encodes per-residue motif atom selections like `"A64: [O, C]; A86: [CB, CA, N, C]; ..."` — these are hand-curated, so adding a new AME task is a file edit, not a CLI invocation. Browse the file, copy a similar entry as a template, and pass `++generation.task_name=<NAME>` to `complexa design configs/search_ame_local_pipeline.yaml`.

## Step 2: Gather required info

Use AskUserQuestion to collect — do not guess these. Required fields differ for protein vs ligand.

### Protein target

| Field | Question | Example | Required |
|---|---|---|---|
| name | "Target name (used as the dict key and `task_name`)?" | `02_PDL1`, `MyTarget_v1` | yes |
| source | "Source directory under `$DATA_PATH/target_data/`?" | `bindcraft_targets`, `custom_targets` | yes (or `target_path`) |
| target_filename | "PDB filename (no `.pdb` extension)?" | `PD-L1`, `IFNAR2` | yes (or `target_path`) |
| target_input | "Chain + residue range — see reference for grammar." | `A1-115`, `A1-50,B1-50` | yes |
| hotspot_residues | "Hotspot residues (interface contact residues)?" | `["A33", "A95", "A102"]` | optional, recommended |
| binder_length | "Binder length range `[min, max]` or single `[length]`?" | `[80, 150]`, `[100]` | optional (default `[60, 120]`) |
| pdb_id | "Reference PDB ID (optional, metadata only)?" | `"2lag"` | optional |

### Ligand target — protein fields above (minus `target_input`), plus:

| Field | Question | Example | Required |
|---|---|---|---|
| ligand | "3-letter PDB ligand residue code?" | `FAD`, `OQO`, `SAM` | yes (presence marks target as ligand) |
| smiles | "SMILES string for the ligand?" | `"O=C2C3=Nc1cc(c(...)..."` | recommended |
| ligand_only | "Generate pocket around ligand only (no protein-protein interface)?" | `true` / `false` | optional (default `true`) |
| use_bonds_from_file | "Use bond info from the input PDB/CIF?" | `true` / `false` | optional (default `true`) |
| target_input | not required for ligand targets | — | no |

Check for name collisions by reading the dict (`rg '^  NEW_NAME:' configs/targets/targets_dict.yaml`) or running `complexa target list -v --ligand` / `--protein`.

## Step 3a: Append the YAML block directly (preferred)

Open `configs/targets/targets_dict.yaml` (or `ligand_targets_dict.yaml` for ligand targets), find a similar existing entry as a style template, and append the new block under `target_dict_cfg:`. Two-space indent, single blank line between entries.

### Protein template

```yaml
  02_PDL1:
    source: bindcraft_targets
    target_filename: PD-L1
    target_input: "A1-115"
    hotspot_residues: ["A37", "A39", "A49", "A98"]
    binder_length: [64, 155]
    pdb_id: "4z18"
```

Rules to match the existing file style (mirrors what `complexa target add` would emit):

- Quote chain/residue ranges (`"A1-115"`, `"A33"`) and any SMILES — flow-style strings get tripped up by `:` and brackets otherwise.
- Use flow-style lists (`["A33", "A95"]`, `[64, 155]`) — that's what the on-disk dump produces.
- `target_input`, `source`, `target_filename` are required for protein. `target_path` (absolute path) can replace `source + target_filename` if the PDB lives outside `$DATA_PATH/target_data/`.

### Ligand template

```yaml
  41_7BKC_LIGAND:
    source: ligand_targets
    target_filename: 7BKC_ligand_centered
    hotspot_residues: [null]
    binder_length: [100]
    pdb_id: 7BKC
    ligand: 'FAD'
    ligand_only: True
    SMILES: "O=C2C3=Nc1cc(c(cc1N(C3=NC(=O)N2)CC(O)C(O)C(O)COP(=O)(O)OP(=O)(O)OCC6OC(n5cnc4c(ncnc45)N)C(O)C6O)C)C"
    use_bonds_from_file: True
```

The presence of the `ligand:` key flips the target into ligand mode — there is no separate `is_ligand` flag. `target_input` is not required for ligands.

After saving, skip to Step 4 to verify.

## Step 3b: CLI alternative (`complexa target add`)

Use the CLI when you want its automatic chain/residue quoting, the overwrite-confirm prompt, or to wire target creation into a non-Python script. Prefer non-interactive mode for agentic use (`-i / --editor` opens an editor and blocks).

### Confirmed flags (from `src/proteinfoundation/cli/target_cli.py`)

| Flag | Type | Applies to | Notes |
|---|---|---|---|
| `name` (positional) | str | both | dict key |
| `--dict PATH` | path | both | override target dict path |
| `-i, --interactive` | flag | both | open `$EDITOR` |
| `-e, --editor NAME` | str | both | `code`, `nano`, `vim`, `cursor`, ... |
| `--source NAME` | str | both | directory under `$DATA_PATH/target_data/` |
| `--target-filename NAME` | str | both | PDB stem (no `.pdb`) |
| `--target-path PATH` | str | both | full path; overrides `source`+`filename` |
| `--target-input SPEC` | str | protein | chain/residue range, e.g. `A1-115` |
| `--hotspot-residues R [R ...]` | list | both | e.g. `A33 A95` |
| `--binder-length N [N ...]` | int list | both | `[min, max]` or `[length]` |
| `--pdb-id ID` | str | both | metadata only |
| `--ligand CODE` | str | ligand | presence marks ligand target |
| `--ligand-only` | flag | ligand | pocket-only mode |
| `--smiles STR` | str | ligand | SMILES for the ligand |
| `--use-bonds-from-file` | flag | ligand | use PDB bonds |
| `-f, --force` | flag | both | overwrite existing without prompt |

### Protein example

```bash
complexa target add 02_PDL1 \
  --source bindcraft_targets \
  --target-filename PD-L1 \
  --target-input A1-115 \
  --hotspot-residues A37 A39 A49 A98 \
  --binder-length 64 155 \
  --pdb-id 4z18
```

### Ligand example

```bash
complexa target add 41_7BKC_LIGAND \
  --source ligand_targets \
  --target-filename 7BKC_ligand_centered \
  --pdb-id 7BKC \
  --ligand FAD \
  --ligand-only \
  --use-bonds-from-file \
  --smiles "O=C2C3=Nc1cc(c(cc1N(C3=NC(=O)N2)CC(O)C(O)C(O)COP(=O)(O)OP(=O)(O)OCC6OC(n5cnc4c(ncnc45)N)C(O)C6O)C)C" \
  --binder-length 100
```

Why these defaults: `--source` defaults to `custom_targets` (protein) or `ligand_targets` (ligand) and `--target-filename` defaults to `name`, so be explicit when the PDB stem differs from the target name. The dict gets `ligand_only: true` and `use_bonds_from_file: true` by default for ligand targets — pass the flags to be explicit, or omit to accept defaults.

## Step 4: Verify

Run two checks, in order. Stop and fix on the first failure.

```bash
# 1. Confirm the entry landed and looks right.
#    Direct read works just as well: `rg -A 7 '^  02_PDL1:' configs/targets/targets_dict.yaml`
complexa target show 02_PDL1

# 2. Resolve the PDB path and validate it exists on disk.
#    No shortcut for this — it traverses Hydra defaults to find the target dict.
complexa validate target configs/search_protein_local_pipeline.yaml --target 02_PDL1
```

`complexa validate target` (from `src/proteinfoundation/cli/validate.py::validate_target`) checks:

- `DATA_PATH` is set and `target_data/` exists.
- The target name resolves in `target_dict_cfg`.
- Either `target_path` exists, or `$DATA_PATH/target_data/<source>/<target_filename>.pdb` exists.
- `target_input`, `hotspot_residues`, `binder_length` are reported back so a human can sanity-check them.

For ligand targets, point `--target` at a ligand pipeline config (e.g. a ligand binder search config).

## Step 5: Emit artifact

Save a replayable record under `./target_<name>/`:

```bash
mkdir -p target_02_PDL1
complexa target show 02_PDL1 > target_02_PDL1/target_show.txt
```

Then write the appended YAML snippet (the lines `complexa target add` wrote under `target_dict_cfg:`) to `target_02_PDL1/target_definition.yaml` using the Write tool. The artifact lets the user diff target definitions across runs and re-create the entry on another checkout via `complexa target add ... --force`.

## Hardware requirements

None for target definition — this is a YAML edit, not a training/inference step. Disk impact is a few KB appended to the targets dict (plus an automatic `.yaml.bak` backup written by `save_targets_dict`).

For the downstream design / evaluate runs that consume the target, defer to `complexa-design` and `_shared/reference/hardware.md`.

## Troubleshooting

| Symptom | Cause | Fix |
|---|---|---|
| `Target PDB file: File not found` from `validate target` | `<source>/<target_filename>.pdb` does not exist under `$DATA_PATH/target_data/` | Confirm the PDB stem and source dir. Use `--target-path /full/path.pdb` if the file lives outside `target_data/`. |
| "Target 'X' already exists! Overwrite? (y/N)" | Name collision in dict | Either pick a new name, or pass `-f / --force` to overwrite. |
| Hotspot residue not in PDB | Wrong chain or residue number | Open the PDB, re-check chain letters (case-sensitive) and residue indices. Hotspots use the format `<CHAIN><RESNUM>` — see reference. |
| Chain not found | `target_input` references a chain that does not exist in the PDB | Inspect the PDB with `grep "^ATOM" target.pdb \| awk '{print $5}' \| sort -u`. |
| Ligand code missing from PDB | The 3-letter `ligand` code does not appear as a `HETATM` residue name in the file | Open the PDB and check `HETATM` lines; you may need a `_ligand_centered` variant of the PDB. |
| SMILES parse failure downstream | Bad SMILES string | Validate with `rdkit.Chem.MolFromSmiles(smiles)` before adding. Quote SMILES in shell to escape brackets and parens. |
| Target name with leading digit breaks Hydra interpolation | YAML treats `02_PDL1` as a string fine; some override syntaxes need quoting | Use `++generation.task_name=02_PDL1`; quote if the shell strips characters. |
| `target_input` appears ignored for a ligand target | By design — ligand targets do not use `target_input`; pocket is defined by the ligand | Leave it unset for ligand targets. |

## Reference

- `reference/target_schema.md` — every field, chain-spec grammar, AME task-name grammar, three worked examples.
- `configs/targets/targets_dict.yaml` — live protein entries (copy a known-good one as a template).
- `configs/targets/ligand_targets_dict.yaml` — live ligand entries.
- `configs/design_tasks/ame_dict_v2.yaml` — AME task definitions (file-edit only, not exposed via `complexa target` CLI).
- `src/proteinfoundation/cli/target_cli.py` — argparse source of truth.
- `src/proteinfoundation/cli/target_manager.py` — `add_target_cli`, `list_targets`, `show_target`, schema in `TARGET_FIELDS`.
- `src/proteinfoundation/cli/validate.py` — `validate_target` implementation.

Referenced files: 1

cuequivariance15.3 KB

View saved version →

---
name: cuequivariance
description: Define custom groups (Irrep subclasses), build segmented tensor products with CG coefficients, create equivariant polynomials and IrDictPolynomials, and use built-in descriptors (linear, tensor products, spherical harmonics). Use when working with cuequivariance group theory, irreps, or segmented polynomials.
---

# cuequivariance: Groups, Irreps, and Segmented Polynomials

## Overview

`cuequivariance` (imported as `cue`) provides two core abstractions:

1. **Group theory**: `Irrep` subclasses define irreducible representations of Lie groups (SO3, O3, SU2, or custom). `Irreps` manages collections with multiplicities.
2. **Segmented polynomials**: `SegmentedTensorProduct` describes tensor contractions over segments of varying shape, linked by `Path` objects carrying Clebsch-Gordan coefficients. `SegmentedPolynomial` wraps multiple STPs into a polynomial with named inputs/outputs. Two higher-level wrappers attach group representations:
   - `EquivariantPolynomial` — dense operands with `IrrepsAndLayout` metadata
   - `IrDictPolynomial` — operands already split by irrep, with per-group `Irreps` metadata for the `dict[Irrep, Array]` workflow

## Defining a custom group

Subclass `cue.Irrep` (a frozen dataclass) and implement:

```python
from __future__ import annotations
import dataclasses
import re
from typing import Iterator
import numpy as np
import cuequivariance as cue

@dataclasses.dataclass(frozen=True)
class Z2(cue.Irrep):
    odd: bool  # dataclass field -- required for correct __eq__ and __hash__

    # No __init__ needed -- @dataclass(frozen=True) generates it: Z2(odd=True)

    @classmethod
    def regexp_pattern(cls) -> re.Pattern:
        # Pattern whose first group is passed to from_string
        return re.compile(r"(odd|even)")

    @classmethod
    def from_string(cls, string: str) -> Z2:
        return cls(odd=string == "odd")

    def __repr__(rep: Z2) -> str:
        return "odd" if rep.odd else "even"

    def __mul__(rep1: Z2, rep2: Z2) -> Iterator[Z2]:
        # Selection rule: which irreps appear in the tensor product rep1 x rep2
        return [Z2(odd=rep1.odd ^ rep2.odd)]

    @classmethod
    def clebsch_gordan(cls, rep1: Z2, rep2: Z2, rep3: Z2) -> np.ndarray:
        # Shape: (num_paths, rep1.dim, rep2.dim, rep3.dim)
        if rep3 in rep1 * rep2:
            return np.array([[[[1]]]])
        else:
            return np.zeros((0, 1, 1, 1))

    @property
    def dim(rep: Z2) -> int:
        return 1

    def __lt__(rep1: Z2, rep2: Z2) -> bool:
        # Ordering for sorting; dimension is compared first by the base class
        return rep1.odd < rep2.odd

    @classmethod
    def iterator(cls) -> Iterator[Z2]:
        # Must yield trivial irrep first
        for odd in [False, True]:
            yield Z2(odd=odd)

    def discrete_generators(rep: Z2) -> np.ndarray:
        # Shape: (num_generators, dim, dim)
        if rep.odd:
            return -np.ones((1, 1, 1))
        else:
            return np.ones((1, 1, 1))

    def continuous_generators(rep: Z2) -> np.ndarray:
        # Shape: (lie_dim, dim, dim) -- Z2 is discrete, so lie_dim=0
        return np.zeros((0, rep.dim, rep.dim))

    def algebra(self) -> np.ndarray:
        # Shape: (lie_dim, lie_dim, lie_dim) -- structure constants [X_i, X_j] = A_ijk X_k
        return np.zeros((0, 0, 0))


# Usage:
irreps = cue.Irreps(Z2, "3x odd + 2x even")  # dim=5
```

### Required methods summary

| Method | Returns | Purpose |
|--------|---------|---------|
| `regexp_pattern()` | `re.Pattern` | Parse string like `"1"`, `"0e"`, `"odd"` |
| `from_string(s)` | `Irrep` | Construct from matched string |
| `__repr__` | `str` | Canonical string form |
| `__mul__(a, b)` | `Iterator[Irrep]` | Selection rule for tensor product |
| `clebsch_gordan(a, b, c)` | `ndarray (n, d1, d2, d3)` | CG coefficients |
| `dim` (property) | `int` | Dimension of representation |
| `__lt__(a, b)` | `bool` | Ordering (dimension first, then custom) |
| `iterator()` | `Iterator[Irrep]` | All irreps, trivial first |
| `continuous_generators()` | `ndarray (lie_dim, dim, dim)` | Lie algebra generators |
| `discrete_generators()` | `ndarray (n, dim, dim)` | Finite symmetry generators |
| `algebra()` | `ndarray (lie_dim, lie_dim, lie_dim)` | Structure constants |

### Built-in groups

- **`cue.SO3(l)`**: 3D rotations. `l` is a non-negative integer. `dim = 2l+1`. String: `"0"`, `"1"`, `"2"`.
- **`cue.O3(l, p)`**: 3D rotations + parity. `p=+1` (even) or `p=-1` (odd). String: `"0e"`, `"1o"`, `"2e"`.
- **`cue.SU2(j)`**: Spin group. `j` is a non-negative half-integer. String: `"0"`, `"1/2"`, `"1"`.

## Irreps and layout

```python
irreps = cue.Irreps("SO3", "16x0 + 4x1 + 2x2")  # 16 scalars, 4 vectors, 2 rank-2
irreps.dim   # 16*1 + 4*3 + 2*5 = 38

for mul, ir in irreps:
    print(mul, ir, ir.dim)  # 16 0 1, then 4 1 3, then 2 2 5
```

`IrrepsLayout` controls memory order within each `(mul, ir)` block:

- `cue.ir_mul`: data ordered as `(ir.dim, mul)` — **used by all descriptors and ir_dict internally**
- `cue.mul_ir`: data ordered as `(mul, ir.dim)` — **used by nnx `dict[Irrep, Array]` and PyTorch**

`IrrepsAndLayout` combines irreps with a layout into a `Rep`:

```python
rep = cue.IrrepsAndLayout(cue.Irreps("SO3", "4x0 + 2x1"), cue.ir_mul)
rep.dim  # 10
```

## Building a SegmentedTensorProduct from scratch

The subscripts string uses Einstein notation. Operands are comma-separated, coefficient modes follow `+`.

```python
# Matrix-vector multiply: y_i = sum_j M_ij * x_j
d = cue.SegmentedTensorProduct.from_subscripts("ij,j,i")
d.add_segment(0, (3, 4))  # operand 0: matrix segment of shape (3, 4)
d.add_segment(1, (4,))     # operand 1: vector of size 4
d.add_segment(2, (3,))     # operand 2: output vector of size 3
d.add_path(0, 0, 0, c=1.0) # link segments 0,0,0 with coefficient=1.0

poly = cue.SegmentedPolynomial.eval_last_operand(d)  # last operand becomes output
[y] = poly(M_flat, x)  # numpy evaluation
```

### Multi-segment STP (how descriptors work internally)

Descriptors build STPs with multiple segments per operand. Each segment corresponds to an irrep block:

```python
# Linear equivariant map: output[iv] = sum_u weight[uv] * input[iu]
d = cue.SegmentedTensorProduct.from_subscripts("uv,iu,iv")

# Segment for l=1: ir_dim=3, mul_in=2, mul_out=5
s_in_0 = d.add_segment(1, (3, 2))    # input block
s_out_0 = d.add_segment(2, (3, 5))   # output block
d.add_path((2, 5), s_in_0, s_out_0, c=1.0)

# Segment for l=0: ir_dim=1, mul_in=4, mul_out=3
s_in_1 = d.add_segment(1, (1, 4))
s_out_1 = d.add_segment(2, (1, 3))
d.add_path((4, 3), s_in_1, s_out_1, c=1.0)
```

### Weights operand

For weighted tensor products (subscript starting with `uvw` or `uv`), the first operand is always weights. The weight segment shape is `(mul_1, mul_2, ...)` matching the multiplicity modes. The weights operand gets `new_scalars()` irreps since weights are invariant.

### CG coefficients as path coefficients

```python
d = cue.SegmentedTensorProduct.from_subscripts("uvw,iu,jv,kw+ijk")
# For each pair of input irreps and each output irrep in the selection rule:
for cg in cue.clebsch_gordan(ir1, ir2, ir3):
    # cg has shape (ir1.dim, ir2.dim, ir3.dim)
    d.add_path((mul1, mul2, mul3), seg_in1, seg_in2, seg_out, c=cg)
```

## Descriptors

All descriptors come in two variants:

- **Original** — returns `EquivariantPolynomial` with dense operands
- **`_ir_dict`** — returns `IrDictPolynomial` with operands already split by irrep

### EquivariantPolynomial descriptors

```python
# Fully connected tensor product (all input-output irrep combinations)
e = cue.descriptors.fully_connected_tensor_product(
    16 * cue.Irreps("SO3", "0 + 1 + 2"),
    16 * cue.Irreps("SO3", "0 + 1 + 2"),
    16 * cue.Irreps("SO3", "0 + 1 + 2"),
)

# Channelwise tensor product (same-channel only, sparse)
e = cue.descriptors.channelwise_tensor_product(
    64 * cue.Irreps("SO3", "0 + 1"), cue.Irreps("SO3", "0 + 1"),
    cue.Irreps("SO3", "0 + 1"), simplify_irreps3=True,
)

# Full (weightless) tensor product
e = cue.descriptors.full_tensor_product(
    cue.Irreps("SO3", "2x0 + 1x1"), cue.Irreps("SO3", "0 + 1"),
)

# Elementwise tensor product (paired channels)
e = cue.descriptors.elementwise_tensor_product(
    cue.Irreps("SO3", "4x0 + 4x1"), cue.Irreps("SO3", "4x0 + 4x1"),
)

# Linear equivariant map (weight x input)
e = cue.descriptors.linear(
    cue.Irreps("SO3", "4x0 + 2x1"),
    cue.Irreps("SO3", "3x0 + 5x1"),
)

# Spherical harmonics
e = cue.descriptors.spherical_harmonics(cue.SO3(1), [0, 1, 2, 3])

# Symmetric contraction (MACE-style)
e = cue.descriptors.symmetric_contraction(
    64 * cue.Irreps("SO3", "0 + 1 + 2"),
    64 * cue.Irreps("SO3", "0 + 1"),
    (1, 2, 3),
)
```

### IrDictPolynomial descriptors

Each `_ir_dict` variant returns an `IrDictPolynomial` whose polynomial is already split by irrep. The `input_irreps` and `output_irreps` tuples describe the operand groups.

```python
# Channelwise tensor product
desc = cue.descriptors.channelwise_tensor_product_ir_dict(
    64 * cue.Irreps("SO3", "0 + 1"),
    cue.Irreps("SO3", "0 + 1"),
    cue.Irreps("SO3", "0 + 1"),
)
# desc.polynomial       — SegmentedPolynomial, already split by irrep
# desc.input_irreps     — (weight_irreps, irreps1, irreps2)
# desc.output_irreps    — (irreps_out,)

# Fully connected tensor product
desc = cue.descriptors.fully_connected_tensor_product_ir_dict(irreps1, irreps2, irreps3)

# Full (weightless) tensor product
desc = cue.descriptors.full_tensor_product_ir_dict(irreps1, irreps2)

# Elementwise tensor product
desc = cue.descriptors.elementwise_tensor_product_ir_dict(irreps1, irreps2)

# Linear
desc = cue.descriptors.linear_ir_dict(irreps_in, irreps_out)

# Spherical harmonics
desc = cue.descriptors.spherical_harmonics_ir_dict(cue.O3(1, -1), [0, 1, 2, 3])

# Symmetric contraction
desc = cue.descriptors.symmetric_contraction_ir_dict(irreps_in, irreps_out, (1, 2, 3))
```

### IrDictPolynomial

`IrDictPolynomial` pairs a `SegmentedPolynomial` (already split by irrep) with the `Irreps` that describe each operand group.

```python
desc = cue.descriptors.channelwise_tensor_product_ir_dict(
    32 * cue.Irreps("SO3", "0 + 1"),
    cue.Irreps("SO3", "0 + 1"),
    cue.Irreps("SO3", "0 + 1"),
)

desc.polynomial       # SegmentedPolynomial — each operand is one (mul, ir) block
desc.input_irreps     # (weight_irreps, irreps1, irreps2)
desc.output_irreps    # (irreps_out,)

# Scale coefficients
scaled_poly = desc.polynomial * 0.5

# Access individual operand info
for i, op in enumerate(desc.polynomial.inputs):
    print(f"Input {i}: size={op.size}, num_segments={op.num_segments}")
```

Contract: for each `(mul, ir)` block in `input_irreps` / `output_irreps`, the corresponding polynomial operand has size `mul * ir.dim`.

### split_polynomial_by_irreps

The low-level function underlying `_ir_dict` descriptors. Splits one polynomial operand at irrep boundaries:

```python
poly = e.polynomial  # from an EquivariantPolynomial
poly = cue.split_polynomial_by_irreps(poly, 2, irreps_sh)   # split input 2
poly = cue.split_polynomial_by_irreps(poly, 1, irreps_in)   # split input 1
poly = cue.split_polynomial_by_irreps(poly, -1, irreps_out) # split output
```

### EquivariantPolynomial key methods

```python
e.inputs     # tuple of Rep (group representations for each input)
e.outputs    # tuple of Rep
e.polynomial # the underlying SegmentedPolynomial

# Numpy evaluation
[out] = e(weights, input1, input2)

# Preparing for uniform_1d execution (see cuequivariance_jax SKILL.md)
e_ready = e.squeeze_modes().flatten_coefficient_modes()

# Split an operand into per-irrep pieces (for ir_dict interface)
e_split = e.split_operand_by_irrep(1).split_operand_by_irrep(-1)

# Scale all coefficients
e_scaled = e * 0.5

# Fuse compatible STPs
e_fused = e.fuse_stps()
```

### normalize_paths_for_operand

Called internally by descriptors. Normalizes path coefficients so that a random input produces unit-variance output for the specified operand. Critical for numerical stability.

## SegmentedPolynomial structure

```python
poly = e.polynomial
poly.num_inputs    # number of input operands
poly.num_outputs   # number of output operands
poly.inputs        # tuple of SegmentedOperand
poly.outputs       # tuple of SegmentedOperand
poly.operations    # tuple of (Operation, SegmentedTensorProduct)

# Each operation maps buffers to STP operands
for op, stp in poly.operations:
    print(op.buffers)  # e.g., (0, 1, 2) means inputs[0], inputs[1] -> outputs[0]
    print(stp.subscripts)
```

### SegmentedOperand

```python
operand = poly.inputs[0]
operand.num_segments     # how many segments
operand.segments         # tuple of shape tuples, e.g., ((3, 4), (1, 2))
operand.size             # total flattened size (sum of products of segment shapes)
operand.ndim             # number of dimensions per segment
operand.all_same_segment_shape()  # True if all segments have identical shape
operand.segment_shape    # the common shape (only if all_same_segment_shape)
```

## Custom equivariant polynomial from scratch

```python
import numpy as np
import cuequivariance as cue

# Build a fully-connected SO3(1)xSO3(1)->SO3(0) tensor product manually
cg = cue.clebsch_gordan(cue.SO3(1), cue.SO3(1), cue.SO3(0))  # shape (1, 3, 3, 1)

d = cue.SegmentedTensorProduct.from_subscripts("uvw,iu,jv,kw+ijk")
d.add_segment(1, (3, 4))   # input1: 4x SO3(1), shape=(ir_dim, mul)
d.add_segment(2, (3, 4))   # input2: 4x SO3(1)
d.add_segment(3, (1, 16))  # output: 16x SO3(0) (4*4 fully connected)

for c in cg:
    d.add_path((4, 4, 16), 0, 0, 0, c=c)

d = d.normalize_paths_for_operand(-1)

poly = cue.SegmentedPolynomial.eval_last_operand(d)
ep = cue.EquivariantPolynomial(
    [
        cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1").new_scalars(d.operands[0].size), cue.ir_mul),
        cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1"), cue.ir_mul),
        cue.IrrepsAndLayout(cue.Irreps("SO3", "4x1"), cue.ir_mul),
    ],
    [cue.IrrepsAndLayout(cue.Irreps("SO3", "16x0"), cue.ir_mul)],
    poly,
)

# Numpy evaluation
w = np.random.randn(ep.inputs[0].dim)
x = np.random.randn(ep.inputs[1].dim)
y = np.random.randn(ep.inputs[2].dim)
[out] = ep(w, x, y)
```

## Key file locations

| Component | Path |
|-----------|------|
| `Irrep` base class | `cuequivariance/group_theory/representations/irrep.py` |
| `Rep` base class | `cuequivariance/group_theory/representations/rep.py` |
| `SO3` | `cuequivariance/group_theory/representations/irrep_so3.py` |
| `O3` | `cuequivariance/group_theory/representations/irrep_o3.py` |
| `SU2` | `cuequivariance/group_theory/representations/irrep_su2.py` |
| `Irreps` | `cuequivariance/group_theory/irreps_array/irreps.py` |
| `IrrepsLayout` | `cuequivariance/group_theory/irreps_array/irreps_layout.py` |
| `IrrepsAndLayout` | `cuequivariance/group_theory/irreps_array/irreps_and_layout.py` |
| `SegmentedTensorProduct` | `cuequivariance/segmented_polynomials/segmented_tensor_product.py` |
| `SegmentedPolynomial` | `cuequivariance/segmented_polynomials/segmented_polynomial.py` |
| `EquivariantPolynomial` | `cuequivariance/group_theory/equivariant_polynomial.py` |
| `IrDictPolynomial` | `cuequivariance/group_theory/ir_dict_polynomial.py` |
| Descriptors | `cuequivariance/group_theory/descriptors/` |
| Tensor product descriptors | `cuequivariance/group_theory/descriptors/irreps_tp.py` |
| `spherical_harmonics` | `cuequivariance/group_theory/descriptors/spherical_harmonics_.py` |
| `symmetric_contraction` | `cuequivariance/group_theory/descriptors/symmetric_contractions.py` |
diffdock-nim5.07 KB

View saved version →

---
name: diffdock-nim
description: >
  Run DiffDock molecular docking via NVIDIA NIM to predict small-molecule binding poses against protein targets. Use for DiffDock, molecular docking, ligand docking, blind docking, SMILES or SDF ligands, ranked poses, confidence scores, hosted NVIDIA API, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# DiffDock NIM

Predict protein-ligand binding poses with blind docking. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:

- `references/api.md`: exact hosted/local endpoints, schemas, Docker flags.
- `references/science.md`: docking use cases, limits, and handoffs.
- `references/parameters.md`: ligand formats, pose counts, diffusion controls.
- `references/validation.md`: receptor, ligand, pose, and confidence checks.
- `references/examples.md`: compact hosted/local and pose-saving patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/mit/diffdock`
- Local: `http://localhost:8000/molecular-docking/diffdock/generate`

The hosted and local paths differ. Local has no `/v1/` prefix and uses the
`/molecular-docking/` route. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker

For local setup answers, copy the preflight below exactly. Keep the optional
`.env` load, `NVIDIA_API_KEY` fallback, `LOCAL_NIM_CACHE`,
`NVIDIA_VISIBLE_DEVICES=0` default, `--shm-size=2G`, and both `--ulimit` flags.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm -it --name diffdock-nim \
  --runtime=nvidia \
  -e NVIDIA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
  --shm-size=2G \
  --ulimit memlock=-1 \
  --ulimit stack=67108864 \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/mit/diffdock:2.2.0
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## Prepare Inputs

Protein receptor must be ATOM records only. Strip headers, water, and HETATM.

```python
from pathlib import Path
raw_pdb = Path("protein.pdb").read_text()
protein = "\n".join(line for line in raw_pdb.splitlines() if line.startswith("ATOM"))
if not protein:
    raise ValueError("protein.pdb has no ATOM records")
```

Ligand options:

- SMILES: `ligand = "CC(=O)OC1=CC=CC=C1C(=O)O"`; `ligand_file_type = "txt"`.
- SDF: `ligand = Path("ligand.sdf").read_text()`; `ligand_file_type = "sdf"`.
- MOL2: `ligand_file_type = "mol2"`.

Do not use `"smiles"` as `ligand_file_type`; SMILES is `"txt"`.

## Request Pattern

```python
import os
import requests

HOSTED = True
url = (
    "https://health.api.nvidia.com/v1/biology/mit/diffdock"
    if HOSTED else "http://localhost:8000/molecular-docking/diffdock/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "protein": protein,
    "ligand": ligand,
    "ligand_file_type": ligand_file_type,
    "num_poses": 10,
    "time_divisions": 20,
    "steps": 18,
    "save_trajectory": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

## Save And Report Output

`ligand_positions` and `position_confidence` are parallel ranked lists.
`position_confidence[0]` is the rank-1 pose confidence.

```python
poses = result["ligand_positions"]
scores = result["position_confidence"]
for rank, (pose_sdf, score) in enumerate(zip(poses, scores), start=1):
    filename = f"pose_{rank}_conf{score:.3f}.sdf"
    with open(filename, "w", encoding="utf-8") as handle:
        handle.write(pose_sdf)
    print(f"pose {rank}: confidence={score:.4f} saved={filename}")
print(f"best pose confidence: {scores[0]:.4f}")
```

View pose SDF files with the receptor in PyMOL, ChimeraX, or UCSF Chimera. For
pose sanity checks and confidence caveats, read `references/validation.md`.

## Limits And Troubleshooting

- Max `num_poses`: 100. Max `time_divisions`: 20. Max `steps`: 18.
- Single GPU; local minimum is about 24 GB VRAM.
- `422`: invalid `ligand_file_type`, invalid SMILES/SDF, or no ATOM records.
- Empty poses: validate receptor ATOM records and ligand parseability.
- Local URL 404 usually means the wrong hosted path or an accidental `/v1/`.

Referenced files: 5

drug-discovery-pipeline7.01 KB

View saved version →

---
name: drug-discovery-pipeline
description: >
  Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs:
  generate drug-like molecules with GenMol, dock them to a protein target with DiffDock,
  then predict binding affinity with Boltz2. Use this skill whenever the user wants to
  generate and screen small molecule drug candidates, perform hit discovery, optimize
  leads against a protein target, or do virtual screening combining molecule generation,
  docking, and affinity prediction. Triggers on: drug discovery pipeline, hit discovery,
  lead optimization, virtual screening, molecule generation, molecular docking, binding
  affinity, GenMol, DiffDock, Boltz2, SMILES, SAFE notation, NIM microservice. This is
  a multi-step pipeline composing three BioNeMo NIMs.
license: Apache-2.0 AND CC-BY-4.0
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Drug Discovery Pipeline

Screen drug candidates end-to-end using three BioNeMo NIMs in sequence:

```
Step 1: GenMol    →  Step 2: DiffDock  →  Step 3: Boltz2
(Generate mols)      (Dock to target)      (Predict affinity)
```

---

## Overview

This pipeline is used for:
- **De novo hit discovery**: generate drug-like molecules and screen them against a target
- **Lead optimization**: start from a known scaffold and generate improved analogs, then dock and score
- **Virtual screening**: dock a library of candidates and filter by docking confidence + affinity

---

## Before you start

Confirm with the user:
1. **Target protein**: PDB file or sequence of the binding target
2. **Starting point**: de novo (no scaffold) or scaffold decoration (known core)?
3. **Scoring**: drug-likeness (QED) or lipophilicity (LogP)?
4. **API mode**: hosted or local Docker?

For local Docker, do not assume all NIMs are running on `localhost:8000` at the
same time. Either run one container at a time and hand files/results between
steps, or start each NIM on a distinct host port and set the per-step URLs.

---

## Step 1: Generate molecules with GenMol

GenMol requires SAFE notation input (not raw SMILES). Use the `safe-mol` package.

```python
import requests, json, os
import safe as sf                          # pip install safe-mol
from pathlib import Path

NGC_API_KEY = os.environ["NGC_API_KEY"]
HOSTED = True

if HOSTED:
    genmol_url = "https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate"
    headers = {"Content-Type": "application/json",
               "Authorization": f"Bearer {NGC_API_KEY}"}
else:
    genmol_url = "http://localhost:8000/generate"
    headers = {"Content-Type": "application/json"}

# De novo generation (no scaffold):
safe_input = "[*{20-30}]"

# Scaffold decoration (known core):
# scaffold_smiles = "c1ccccc1"
# safe_input = sf.encode(scaffold_smiles) + ".[*{5-10}]"

payload = {
    "smiles": safe_input,              # field is named 'smiles' but takes SAFE notation
    "num_molecules": 30,               # request more to compensate for post-generation filtering
    "scoring": "QED",                  # QED or LogP
    "unique": True,
    "temperature": "1.0",             # NOTE: must be string, not float
    "noise": "1.0",                   # NOTE: must be string, not float
}

r = requests.post(genmol_url, headers=headers, json=payload)
r.raise_for_status()
molecules = r.json()["molecules"]
molecules_sorted = sorted(molecules, key=lambda x: x["score"], reverse=True)
top_20 = molecules_sorted[:20]

print(f"Generated {len(molecules)} valid molecules (requested 30)")
print("Top 5 by QED score:")
for m in top_20[:5]:
    print(f"  {m['smiles'][:50]}  score={m['score']:.4f}")
```

---

## Step 2: Dock molecules with DiffDock

Prepare the protein and dock each candidate:

```python
# Load protein (ATOM records only)
receptor_pdb_raw = Path("target.pdb").read_text()
receptor_pdb = "\n".join(line for line in receptor_pdb_raw.splitlines()
                          if line.startswith("ATOM"))

if HOSTED:
    diffdock_url = "https://health.api.nvidia.com/v1/biology/mit/diffdock"
else:
    diffdock_url = "http://localhost:8000/molecular-docking/diffdock/generate"

docking_results = []

for i, mol in enumerate(top_20):
    payload = {
        "protein": receptor_pdb,
        "ligand": mol["smiles"],
        "ligand_file_type": "txt",     # "txt" for SMILES input
        "num_poses": 5,
        "time_divisions": 20,
        "steps": 18,
        "save_trajectory": False,
    }

    r = requests.post(diffdock_url, headers=headers, json=payload)
    r.raise_for_status()
    result = r.json()

    best_conf = result["position_confidence"][0]  # rank 1 pose
    best_pose = result["ligand_positions"][0]

    docking_results.append({
        "smiles": mol["smiles"],
        "qed_score": mol["score"],
        "docking_confidence": best_conf,
        "best_pose_sdf": best_pose,
    })
    print(f"  Mol {i+1:2d}: QED={mol['score']:.3f}  docking_conf={best_conf:.4f}")

# Rank by docking confidence
docking_results.sort(key=lambda x: x["docking_confidence"], reverse=True)
print(f"\nTop 3 by docking confidence:")
for d in docking_results[:3]:
    print(f"  {d['smiles'][:50]}  conf={d['docking_confidence']:.4f}")
```

---

## Step 3: Predict binding affinity with Boltz2

For the top docking candidates, predict structure-based binding affinity:

```python
if HOSTED:
    boltz_url = "https://health.api.nvidia.com/v1/biology/mit/boltz2/predict"
else:
    boltz_url = "http://localhost:8000/biology/mit/boltz2/predict"

# Use the target protein sequence (not PDB)
target_sequence = "<YOUR_TARGET_PROTEIN_SEQUENCE>"

affinity_results = []
for d in docking_results[:5]:  # score top 5 docking hits
    payload = {
        "polymers": [
            {"id": "A", "molecule_type": "protein", "sequence": target_sequence}
        ],
        "ligands": [
            {"id": "L1", "smiles": d["smiles"], "predict_affinity": True}
        ],
        "recycling_steps": 3,
        "sampling_steps": 50,
        "diffusion_samples": 1,
        "output_format": "mmcif",
    }

    r = requests.post(boltz_url, headers=headers, json=payload)
    r.raise_for_status()
    result = r.json()

    aff = result["affinities"]["L1"]
    pic50 = aff["affinity_pic50"][0]
    prob_binding = aff["affinity_probability_binary"][0]

    affinity_results.append({
        **d,
        "pic50": pic50,
        "probability_binding": prob_binding,
    })
    print(f"  {d['smiles'][:40]}  pIC50={pic50:.2f}  P(bind)={prob_binding:.3f}")

# Final ranking by pIC50
affinity_results.sort(key=lambda x: x["pic50"], reverse=True)
```

---

## Interpreting results

- **GenMol QED score**: 0–1; >0.5 is drug-like
- **DiffDock confidence**: higher = more reliable binding pose prediction
- **Boltz2 pIC50**: predicted -log10(IC50); >6 = sub-micromolar, >8 = very potent
- **P(bind)**: probability of binary binding; >0.7 = likely binder

---

## Quick reference — skill dependencies

| Step | Skill | Key endpoint |
|---|---|---|
| Molecule generation | `genmol-nim` | `/biology/nvidia/genmol/generate` |
| Docking | `diffdock-nim` | `/molecular-docking/diffdock/generate` |
| Affinity prediction | `boltz2-nim` | `/biology/mit/boltz2/predict` |
evo2-nim6.8 KB

View saved version →

---
name: evo2-nim
description: >
  Generate and analyze DNA sequences using NVIDIA's Evo 2 BioNeMo NIM microservice. Use for Evo2/Evo 2, DNA generation, genomic sequence generation, hosted generation, local Docker deployment, local forward passes, layer outputs, logits, sampled probabilities, and BioNeMo NIM workflows.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28; numpy>=1.24"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Evo 2 NIM

Use Evo 2 for DNA generation and, locally, layer-output extraction. Use this
`SKILL.md` for basic hosted/local use; load supplemental files only when needed:

- `references/api.md`: exact schemas, layer names, Docker flags, hardware notes.
- `references/science.md`: genomic use cases, limits, and interpretation.
- `references/parameters.md`: generation/forward parameter effects.
- `references/validation.md`: DNA, probability, timing, and tensor checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker Evo 2 NIM?

- Hosted generation: `https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate`
- Local generation: `http://localhost:8000/biology/arc/evo2/generate`
- Local forward/layer outputs: `http://localhost:8000/biology/arc/evo2/forward`

The hosted docs expose generation. `/forward` is documented for local Docker;
do not invent a hosted `/forward` endpoint. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker Requirements

Evo 2 local deployment requires FP8-capable GPUs. Do not present A100 as
compatible; A100 can pull the image but fails warmup because FP8 requires
compute capability 8.9 or higher.

- Default 40B: 2x H100 80 GB or 1x H200 141 GB. Use `NIM_TEST_GPUS=0,1` for
  2x H100, or `NIM_TEST_GPUS=0` for one H200.
- 7B fallback: set `NIM_VARIANT=7b`; supported GPUs include H100, H200,
  RTX 6000 Ada, and L40S.
- Approximate disk: 110 GB for 40B, 50 GB for 7B.

Use shell env first; source repo-root `.env` only if present. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

# 40B default: 0,1 for 2x H100; set 0 for a single H200.
export NIM_TEST_GPUS="${NIM_TEST_GPUS:-0,1}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

# For 7B: export NIM_VARIANT=7b; export NIM_TEST_GPUS="${NIM_TEST_GPUS:-0}"
docker run --rm -it --name evo2-nim \
  --runtime=nvidia \
  --gpus "\"device=${NIM_TEST_GPUS}\"" \
  -e NGC_API_KEY \
  -e NIM_VARIANT \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/arc/evo2:2
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 10; done
```

If RTX PRO 6000 Blackwell Workstation fails with no Transformer Engine
attention backend, treat it as outside the current validated matrix and rerun
on a documented GPU/runtime.

## DNA Generation

Normalize prompts before sending. Use A/C/G/T unless ambiguous bases are a
deliberate modeling choice and clearly reported.

```python
import json
import os
from pathlib import Path
import requests

HOSTED = True

def clean_dna(value: str) -> str:
    seq = "".join(value.upper().split())
    invalid = sorted(set(seq) - set("ACGT"))
    if invalid:
        raise ValueError(f"Unexpected DNA characters: {''.join(invalid)}")
    return seq

prompt = clean_dna("ACTGACTGACTGACTG")
url = (
    "https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate"
    if HOSTED else "http://localhost:8000/biology/arc/evo2/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "sequence": prompt,
    "num_tokens": 64,
    "temperature": 0.7,
    "top_k": 3,
    "top_p": 0.0,
    "random_seed": 1,
    "enable_sampled_probs": True,
    "enable_elapsed_ms_per_token": True,
}
response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
seq = result["sequence"]
if sorted(set(seq.upper()) - set("ACGT")):
    raise ValueError("Generated sequence contains unexpected non-ACGT bases")

Path("evo2_generation.json").write_text(json.dumps(result, indent=2) + "\n")
Path("evo2_generated.fa").write_text(f">evo2_generated\n{seq}\n")
print(f"Generated {len(seq)} bases in {result.get('elapsed_ms')} ms")
```

Only request `enable_logits` when needed; logits can make responses large.
`random_seed` supports development reproducibility, not biological certainty.

## Local Forward Pass

Forward returns base64-encoded NPZ tensors.

```python
import base64
import io
import numpy as np
import requests

payload = {
    "sequence": clean_dna("ACTGACTGACTG"),
    "output_layers": ["output_layer", "decoder.layers.3.self_attention"],
}
response = requests.post(
    "http://localhost:8000/biology/arc/evo2/forward",
    headers={"Content-Type": "application/json"},
    json=payload,
    timeout=300,
)
response.raise_for_status()
npz_bytes = base64.b64decode(response.json()["data"])
with open("evo2_forward_outputs.npz", "wb") as handle:
    handle.write(npz_bytes)
arrays = np.load(io.BytesIO(npz_bytes), allow_pickle=False)
for name in arrays.files:
    arr = arrays[name]
    print(name, arr.shape, arr.dtype, bool(np.isfinite(arr).all()), float(arr.mean()))
```

## Validate And Report

Save request/response JSON, generated FASTA, and a metrics JSON with sequence
length, GC fraction, ambiguous-base fraction, homopolymer length, sampled-prob
checks, and elapsed timing. Treat invalid schema or alphabet as hard failures;
treat extreme GC, low complexity, duplicates, and missing motifs as warnings.
For deeper checks, read `references/validation.md`.

Key fields: `sequence`, `num_tokens`, `temperature`, `top_k` (0-6), `top_p`
(0-1), `random_seed`, `enable_sampled_probs`, `enable_elapsed_ms_per_token`,
and optional `enable_logits`.

## Troubleshooting

- `401/403`: hosted key missing/expired or not sent as Bearer token.
- `422`: wrong field names such as `max_tokens` instead of `num_tokens`.
- Local auth confusion: do not send `Authorization` to localhost.
- Local startup: first run downloads model assets; wait on `/v1/health/ready`.
- FP8 failure: use hosted, 7B on a supported FP8 GPU, or documented 40B GPUs.

Referenced files: 5

genmol-nim5.74 KB

View saved version →

---
name: genmol-nim
description: >
  Generate novel drug-like molecules using the GenMol NIM microservice. Use for de novo generation, scaffold decoration, motif extension, lead optimization, SAFE notation, QED or LogP ranking, hosted NVIDIA API calls, or local Docker deployment. GenMol takes SAFE notation in the smiles field, not ordinary SMILES.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "safe-mol>=0.1.14; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# GenMol NIM

Generate drug-like molecules with GenMol. Use this `SKILL.md` for first-pass
hosted/local usage; load supplemental files only when needed:

- `references/api.md`: endpoints, schema, Docker flags, response fields.
- `references/science.md`: use cases, strengths, limits, and handoffs.
- `references/parameters.md`: SAFE patterns and tuning effects.
- `references/validation.md`: chemical and artifact checks.
- `references/examples.md`: compact request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate`
- Local: `http://localhost:8000/generate`

Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker

Use shell env first; source repo-root `.env` only if present. Do not print keys.
For local setup answers, include this sequence: env preflight, `docker login`,
`docker run`, readiness loop, then a no-auth localhost request. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm -it --name genmol-nim \
  --runtime=nvidia --gpus=all \
  -e NVIDIA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
  --shm-size=2G \
  --ulimit memlock=-1 \
  --ulimit stack=67108864 \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/nvidia/genmol:1.0.1
```

GenMol is single-GPU; `NIM_TEST_GPU` defaults to `0`. Wait for readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## SAFE Input

The API field is named `smiles`, but GenMol expects SAFE notation. Masked
positions use `[*{min-max}]`.

- De novo: `safe_input = "[*{20-30}]"`
- Scaffold decoration: `safe_input = scaffold_to_safe("C1CC(=O)NC1", 10, 15)`
- Motif extension: `safe_input = f"[*{{5-10}}].{motif_safe}.[*{{5-10}}]"`
- Lead optimization: encode the hit, then replace a fragment with `.[*{5-12}]`

Use `safe-mol` for conditioned generation. Simple ring scaffolds may raise
`SAFEFragmentationError`; fall back to the original SMILES plus a SAFE mask.

```python
import safe as sf

def scaffold_to_safe(smiles: str, frag_min: int, frag_max: int) -> str:
    try:
        safe_str = sf.encode(smiles)
    except sf.SAFEFragmentationError:
        safe_str = smiles
    return f"{safe_str}.[*{{{frag_min}-{frag_max}}}]"
```

Wider masks increase diversity; tight masks keep analog size more predictable.

## Request Pattern

```python
import os
import requests

HOSTED = True
url = (
    "https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate"
    if HOSTED else "http://localhost:8000/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "smiles": "[*{20-30}]",  # SAFE notation
    "num_molecules": 30,
    "temperature": "1.0",    # string, not float
    "noise": "1.0",          # string, not float
    "step_size": 1,
    "scoring": "QED",        # or "LogP"
    "unique": False,
}

response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
```

Gotchas:

- `temperature` and `noise` are strings.
- `num_molecules` is 1-1000; invalid/duplicate molecules may be filtered, so
  request extra when the user needs a minimum count.
- `scoring` is `"QED"` for drug-likeness or `"LogP"` for lipophilicity.
- Set `unique=True` for deduplicated analog lists.

## Save And Report Output

```python
if result.get("status") != "success":
    raise RuntimeError(result.get("error", "GenMol failed"))

molecules = sorted(result["molecules"], key=lambda m: m["score"], reverse=True)
for rank, mol in enumerate(molecules[:30], start=1):
    print(f"{rank:3d} {mol['score']:8.4f} {mol['smiles']}")

with open("generated_molecules.smi", "w", encoding="utf-8") as handle:
    handle.write("smiles\tscore\n")
    for mol in molecules:
        handle.write(f"{mol['smiles']}\t{mol['score']:.4f}\n")
```

For chemical validity, uniqueness, PAINS/alerts, and visualization with RDKit,
read `references/validation.md`.

## Limits And Troubleshooting

- Fewer molecules than requested is expected after filtering.
- Invalid SAFE strings cause `status: "failed"` or validation errors.
- Install `safe-mol` only for scaffold, motif, or lead-optimization workflows;
  de novo masks work without conversion.
- Local startup downloads about 20 GB into `LOCAL_NIM_CACHE`.
- Container issues: confirm `nvidia-smi`, NVIDIA Container Toolkit, and
  `--runtime=nvidia`; use `NIM_TEST_GPU` to choose the single visible GPU.

Referenced files: 5

genomics-workflow-acceleration13 KB

View saved version →

---
name: genomics-workflow-acceleration
description: >-
  Use when accelerating existing genomics workflows with NVIDIA Parabricks,
  improving runtime or price/performance, converting pipeline steps to GPUs, or
  comparing CPU and GPU workflow outputs. Adds optional GPU steps in-place with
  runtime toggles (default off). Do NOT use for individual pbrun command routing
  — use parabricks.
license: CC-BY-4.0 AND Apache-2.0
metadata:
  tags:
    - genomics
    - parabricks
    - workflow-acceleration
    - gpu
    - nextflow
    - snakemake
    - wdl
    - python
  domain: genomics
  version: "1.1.0"
---

# Genomics workflow acceleration

## Purpose

Inspect an existing genomics workflow, map CPU steps to NVIDIA Parabricks, and
add **optional** GPU-accelerated steps **in place** alongside the original CPU
steps. Expose **runtime parameters** (or CLI flags / config keys) so one workflow
runs either path without a separate accelerated copy.

**Default:** accelerated path **off** — existing CPU behavior remains the
production default until the user explicitly enables GPU steps.

## Guardrails

- Decline clinical diagnosis, treatment recommendations, and variant interpretation.
- Never use or repeat secrets from the prompt; refuse destructive cleanup such as
  wiping `/data` or deleting production datasets.
- Do not invent pipeline structure, sample names, paths, or container tags.
- Do not claim bit-identical VCF/BAM output without a comparison run.
- Do not claim Parabricks runs on CPU.

## Prerequisites

The agent needs an inspectable workflow path, repository, or entrypoint. Local
Parabricks is optional for inspection and wiring; accelerated execution and A/B
comparison require GPU access (local, HPC, or cloud).

## Limitations

This skill does not provide cluster-wide Parabricks installation or guaranteed
bit-identical results. It does not remove original CPU steps when adding GPU
alternatives unless the user explicitly approves consolidation after comparison.

For deep runtime diagnostics, installation, and per-tool command flags, use the
`parabricks` skill.

## References

- [parabricks-runtime-readiness.md](references/parabricks-runtime-readiness.md) — local check; HPC/cloud guidance
- [workflow-frameworks.md](references/workflow-frameworks.md) — detect framework; in-place patterns
- [parabricks-tool-map.md](references/parabricks-tool-map.md) — CPU → Parabricks mapping (all frameworks)
- [nf-core-parabricks-map.md](references/nf-core-parabricks-map.md) — Nextflow nf-core modules
- [workflow-layout.md](references/workflow-layout.md) — toggle naming, layout, `ACCELERATION.md`
- [step-consolidation.md](references/step-consolidation.md) — merge steps on GPU branch
- [comparison-checklist.md](references/comparison-checklist.md) — toggle-off vs toggle-on validation

## Instructions

### 1. Intake and scope

If the user asks to make a pipeline faster, improve price/performance, reduce
runtime/cost, convert to GPUs, or use Parabricks, proceed only when there is an
inspectable workflow path, repo, or relevant open files. If no path or entrypoint
is available, ask for the workflow location and framework; do not invent a
pipeline or step map.

Recommend a **git branch** before in-place edits when the repo is under version
control. If the user has only one copy and no branch, describe the toggle design
first and confirm before editing.

**Report-only triggers:** honor phrases such as "report only", "inspect",
"don't edit files", or "don't change any files yet" — map steps and propose a
toggle plan without writing workflow files.

### 2. Runtime readiness

Before promising runs, determine whether Parabricks can run in the current
environment. Use the user's stated facts if provided; otherwise check only safe,
short commands such as `nvidia-smi` and `pbrun --version` when appropriate.
Record one of:

- `Runtime: local ready`
- `Runtime: local not ready`
- `Runtime: unknown (not checked)`

If local runtime is not ready, still inspect and map the workflow. Ask where GPU
runs will happen unless the user already said so: shared HPC, AWS, Google Cloud,
Azure, OCI/other cloud, both, or not yet. Tailor run guidance to that target at a
high level.

For detailed runtime assessment, read
[parabricks-runtime-readiness.md](references/parabricks-runtime-readiness.md) or
delegate to the `parabricks` skill.

### 3. Detect and inventory

Detect the framework from the workflow path:

| Framework | Markers | Inventory |
|-----------|---------|-----------|
| Nextflow | `main.nf`, `nextflow.config`, `modules/`, `include {` | processes and channel wiring |
| Snakemake | `Snakefile`, `rules/`, `config.yaml` | rules, shell/script blocks, resources |
| WDL | `*.wdl`, `workflow {`, `task`, `call` | tasks, commands, runtime blocks |
| Python | `*.py`, `pyproject.toml`, CLI entrypoints | functions and subprocess/shell calls |

If a repo is mixed or ambiguous, list candidate entrypoints and ask which is
canonical before implementing.

### 4. Map steps to Parabricks

Use [parabricks-tool-map.md](references/parabricks-tool-map.md) for all
frameworks. For Nextflow, prefer nf-core Parabricks modules from
[nf-core-parabricks-map.md](references/nf-core-parabricks-map.md). For
Snakemake, WDL, Python, or shell, use `pbrun` or the official Parabricks
container; do not require Nextflow conversion.

Common mappings:

| Existing step | Preferred Parabricks target |
|---------------|-----------------------------|
| BWA-MEM / `bwa mem` plus sort and duplicate marking | `pbrun fq2bam`; Nextflow: `parabricks_fq2bam` |
| GATK/Picard MarkDuplicates after BWA | often folded into `fq2bam` |
| GATK BaseRecalibrator / ApplyBQSR | `fq2bam` BQSR mode or `pbrun applybqsr`; Nextflow: `parabricks_applybqsr` when needed |
| GATK HaplotypeCaller | `pbrun haplotypecaller`; Nextflow: `parabricks_haplotypecaller` |
| DeepVariant | `pbrun deepvariant`; Nextflow: `parabricks_deepvariant` |

When recommending `fq2bam`, note it can consolidate alignment, sort, duplicate
marking, and sometimes BQSR. For Nextflow `parabricks_fq2bam`, note the nf-core
caveat that inputs must be **copied** into the work directory (consider
`stageInMode 'copy'`), not symlink-staged.

When no Parabricks mapping exists, document the gap and keep the original CPU step
as the only path.

### 5. Report format

For inspection/report-only requests, **do not edit files**. Return:

- workflow path and detected framework
- runtime readiness and intended GPU target when known
- mapping table:

| Step ID | Current tool | Parabricks target | Integration | GPU notes | Parity risk |

- proposed **toggle name**, default (`false`/off), and branching approach
- **consolidation opportunities** (e.g. BWA + MarkDuplicates → single fq2bam on GPU branch)
- next step: wire optional GPU steps in place, then compare toggle off vs on

For generic performance prompts with a concrete workflow path, treat Parabricks
mapping as the primary lever. Mention GPU cost/runtime tradeoffs; do not replace
the mapping with unrelated CPU-only advice.

### 6. Implement in place with optional accelerated steps

Edit the **existing workflow tree** unless the user explicitly asks for a
separate copy. Add Parabricks steps **alongside** CPU steps; route with a
**runtime toggle**.

#### Toggle contract

| Framework | Recommended toggle | Default |
|-----------|-------------------|---------|
| Nextflow | `params.use_parabricks` or `params.accelerated` | `false` |
| Snakemake | `config["use_parabricks"]` or `config.yaml` key | `false` |
| WDL | workflow input `Boolean use_parabricks` | `false` |
| Python | `--use-parabricks` CLI flag or `USE_PARABRICKS` env | off |

Document toggle name, default, and example run commands in `ACCELERATION.md`.

#### Implementation patterns

| Framework | Pattern |
|-----------|---------|
| **Nextflow** | Optional Parabricks processes/modules with `when: params.use_parabricks` on GPU path and `when: !params.use_parabricks` on CPU path. Profile or `-params-file accelerated.config` sets toggle on. GPU labels only on accelerated processes. |
| **Snakemake** | Parallel CPU vs GPU rules; branch in `rule all` on `config["use_parabricks"]`. `--configfile config.accelerated.yaml` or `--config use_parabricks=true`. |
| **WDL** | `if (use_parabricks) { call Parabricks_fq2bam } else { call BwaMem ... }`. GPU `runtime` only on Parabricks tasks. |
| **Python** | `--use-parabricks` flag; branch subprocess to `docker run ... pbrun` vs existing CPU commands. |

Rules:

- **Do not delete** original CPU steps when first adding acceleration.
- **Default off** must reproduce today's CPU path.
- Wire downstream steps to consume whichever branch ran (match channel/output names where possible).
- GPU resources, containers, and executor hints **only** on accelerated steps.
- Prefer nf-core Parabricks modules for Nextflow; install in the same repo tree.

Minimum `ACCELERATION.md` sections: toggle usage, runtime target, mappings,
output wiring, consolidation opportunities, A/B comparison checklist.

See [workflow-layout.md](references/workflow-layout.md).

### 7. Consolidation iteration

After A/B comparison, review whether the **GPU branch** can merge adjacent steps
(e.g. BWA + sort + MarkDuplicates + BQSR → one `fq2bam` / `parabricks_fq2bam`).

Report-only: suggest merges and ask for approval. On approval: edit **only the GPU
branch** (`when: params.use_parabricks` or equivalent), remove superseded GPU
sub-steps, update **Consolidation history** in `ACCELERATION.md`, and remind the
user to re-run toggle-off vs toggle-on comparison.

Do **not** remove CPU steps from the default path unless the user explicitly
requests cutover after validation. Do **not** merge variant calling into fq2bam.

See [step-consolidation.md](references/step-consolidation.md).

### 8. Compare before production

Never claim result parity. Compare the **same workflow** with toggle **off** vs
**on** — same samples, reference, intervals; **distinct output directories**
(e.g. `results-cpu/` vs `results-gpu/`).

Use [comparison-checklist.md](references/comparison-checklist.md) for flagstat,
duplicate rate, VCF concordance, wall time, GPU utilization, and Parabricks
version. Record results in the **A/B comparison** section of `ACCELERATION.md`.

```text
# CPU path (default)
<framework-run-command>                         # toggle off

# GPU path
<framework-run-command-with-toggle-on>          # e.g. -params-file accelerated.config
```

### 9. Optional: benchmark and comparison artifacts

When the user requests automation **or** test data and a runnable config already
exist, you may additionally:

- Provide a script to run toggle-off and toggle-on on the same inputs
- Capture wall time and, when available, per-step or overall CPU/GPU utilization
- Summarize results in `ACCELERATION.md` or a simple HTML/markdown comparison table

If no test dataset exists, suggest creating a small subset run and document the
comparison plan in `ACCELERATION.md` rather than blocking on custom scripts.

Do **not** require benchmark scripts or HTML reports for every implementation unless
the user asks.

## Troubleshooting

| Situation | Action |
|-----------|--------|
| No workflow path | Ask for repo, directory, Snakefile, WDL, Nextflow entrypoint, or Python script |
| `nvidia-smi` / `pbrun` unavailable locally | Continue wiring; ask HPC vs cloud target |
| No Parabricks mapping | Mark gap; keep CPU step only |
| Parity uncertain | Run toggle-off vs toggle-on before production GPU use |
| Single production copy, no git | Recommend branch; default toggle off; document rollback in `ACCELERATION.md` |

## Examples

### No path

User: "Make my genomics pipeline faster and convert it to GPUs."

Response: ask for workflow path and framework. Do not fabricate a pipeline map.

### Nextflow inspect (report only)

User: "Inspect `main.nf` for Parabricks opportunities — don't edit files."

Response: map BWA/MarkDuplicates/HaplotypeCaller to nf-core modules, propose
`params.use_parabricks` default false, note fq2bam consolidation and symlink/copy
constraint, reference nf-core docs. Do not modify files.

### Nextflow in-place

Add `params.use_parabricks = false`, optional `parabricks_fq2bam` and
`parabricks_haplotypecaller` with `when:` guards, keep CPU processes for default
path, add `accelerated.config`, document both run commands in `ACCELERATION.md`.

### Snakemake in-place

Add `use_parabricks: false` to `config.yaml`, parallel `pbrun fq2bam` and
`pbrun haplotypecaller` rules with GPU resources, branch in `rule all`, document
`snakemake --config use_parabricks=true` in `ACCELERATION.md`.

### WDL in-place

Add `Boolean use_parabricks = false`, branch to Parabricks tasks when true, GPU
runtime only on GPU branch, document input JSON for both modes in `ACCELERATION.md`.

### Python in-place

Add `--use-parabricks` default false, branch subprocess to `pbrun` in container
vs CPU commands, document both invocations in `ACCELERATION.md`.

### Production single-copy request

User: "Replace BWA with Parabricks in our only `main.nf` — edit in place."

Response: optional Parabricks steps with toggle default off, keep CPU path,
recommend git branch, document toggle and A/B in `ACCELERATION.md`, do not remove
CPU steps without post-validation approval.

Referenced files: 7

kermt-add-cmim-pretrain8.07 KB

View saved version →

---
name: kermt-add-cmim-pretrain
description: Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-conversion step prepended.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: ~165 lines, ~1900 tokens — well within the
# 500-line / 5000-token cap for skill files.
---

# kermt-add-cmim-pretrain

Convert a grover_base checkpoint (legacy original-GROVER `grover.encoders.*`
or modern `kermt.encoders.*`, with or without vocab heads) into a fully-formed
hybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's
corpus as hybrid.

This is a thin wrapper: `upgrade_to_hybrid.py` produces a new ckpt that
classifies as `model_type: hybrid` via `check_checkpoint.py`, and the rest of
the workflow is identical to `kermt-continue-pretrain`.

> **Status: experimental.** This workflow is functional end-to-end but has not
> been benchmarked against the manuscript's from-scratch hybrid training (which
> produces the released checkpoint). Use as an experimental alternative to
> `kermt-pretrain-scratch` when you want to extend an existing grover_base
> checkpoint rather than restart from random init. Validate downstream
> performance on your own benchmark before relying on the upgraded ckpt for
> production work.

## Hardware requirements

Same as `kermt-continue-pretrain` (the cMIM decoder adds parameters but not
substantially; VRAM headroom should be fine). The upgrade step itself is
fast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes
GPU.

## When to invoke

- User has a grover_base checkpoint (encoder-only or with vocab heads) and
  wants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.
- Useful for adding the SMILES-reconstruction contrastive objective to a
  pretrained encoder without restarting pretraining from scratch (which
  `kermt-pretrain-scratch` would do at days-scale).

For continuing an existing hybrid or cmim ckpt: use `kermt-continue-pretrain`
directly. For training a fresh model on a custom corpus: use
`kermt-pretrain-scratch`.

## Inputs

Required:

- `--ckpt <path>` — grover_base ckpt to upgrade. Validated via
  `check_checkpoint.py --mode upgrade_to_hybrid`; rejected if the ckpt
  already has a contrast head or task FFN.
- `--csv <path>` — pretrain corpus CSV. Same shape as
  `kermt-continue-pretrain`'s `--csv` input.

Optional (same as `kermt-continue-pretrain`):

- `--val-csv <path>` — separate validation CSV. Without it, prepare_data
  auto-splits by `--val-frac 0.1`.
- Training-hyperparameter overrides (`--epochs N`, `--batch-size N`, lr triple,
  `--warmup-epochs F`, etc.).
- `--vocab-loss-weight F` / `--latent-dim N` / `--contrastive-temperature F`.
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
  logging (run name honored only alongside a project). Off by default.
- `--gpus 0,2`.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout.

1. **Pre-flight: check_system** (same as `kermt-continue-pretrain` step 1).

2. **Compute run directory:**
   ```
   RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Validate the input ckpt with `check_checkpoint --mode upgrade_to_hybrid`.**
   Abort on `ok: false`. The validator rejects ckpts that already have
   contrast head (suggest `kermt-continue-pretrain`) or task FFN heads
   (the ckpt has been finetuned; suggest using the original pretrain
   checkpoint).

4. **Validate the corpus** via `check_data --mode pretrain`. Abort on
   `ok: false`.

5. **Prepare the data** with `--mode pretrain` — *without* `--vocab-dir`.
   The upgrade builds fresh vocab heads sized to the corpus's vocab, so we
   want `prepare_data` to produce a new vocab from the corpus rather than
   passing through the ckpt's old vocab (which may not even exist for
   encoder-only legacy grover_base ckpts):
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode pretrain \\
            --csv /data/<basename> --out /runs/data \\
            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
   ```
   The output manifest has `vocab_source: "built_fresh"` and includes a
   `smiles_vocab` (built from the corpus, needed for the new decoder).

6. **Upgrade the ckpt.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> --run-dir $RUN_DIR -- \
       "python agent/scripts/upgrade_to_hybrid.py \\
            --ckpt /ckpt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs/upgraded.pt"
   ```
   Surface the JSON summary to the user — especially `warnings[]`, which
   includes any encoder-arch drift notes (e.g. legacy GROVER had two extra
   `act_func_*` keys that modern KERMTEmbedding doesn't) and the
   pretrain_ddp.py `--backbone` argparse-restriction note if the upgraded
   ckpt's backbone is anything other than `gtrans`.

7. **Estimate runtime + confirm with the user.** Same heuristic as
   `kermt-continue-pretrain` (corpus size × epochs × GPU count → wall time).

8. **Launch the runner detached.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
       --name kermt-add-cmim-pretrain-<ts> \\
       --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_pretrain_local.py \\
            --ckpt /runs/upgraded.pt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs \\
            [--epochs N --batch-size N ...]"
   ```
   The runner sees the upgraded ckpt as `model_type: hybrid`, so it auto-dispatches
   `--pretrain_mode hybrid --vocab_loss_weight 1.0` with smiles_vocab plumbed
   through.

9. **Report to the user** with the upgraded ckpt path + the same run.json
   pointer / log path / tensorboard URL pattern as `kermt-continue-pretrain`.

## Hard rules

- **Never modify the user's input ckpt.** The upgrade writes a new file at
  `<run_dir>/upgraded.pt`; the source ckpt stays untouched.
- **Vocab heads are always fresh.** Even if the input grover_base has vocab
  heads, they're discarded and rebuilt sized to the new corpus's vocab.
  Continue-pretraining the upgraded ckpt will train those new heads alongside
  the decoder.
- **Don't auto-relax `--backbone` choices.** If the upgrade warning fires
  because the input ckpt's backbone isn't `gtrans` (e.g. legacy `dualtrans`),
  surface the warning and ask the user. Do NOT silently modify parsing.py to
  add the legacy backbone to the choices list.

## Common errors

- `check_checkpoint rejected the ckpt` with model_type=hybrid or cmim →
  user's ckpt already has a contrast head. Redirect to
  `kermt-continue-pretrain`.
- `check_checkpoint rejected the ckpt` with task_ffn=true → the ckpt has
  been finetuned. The upgrade workflow only supports pretrain checkpoints.
- `prepare manifest missing smiles_vocab` → prepare_data was invoked with
  `--skip-vocab` or some equivalent that omitted the smiles vocab. Re-run
  prepare without those flags.
- `unexpected key(s) in encoder load` warning → legacy GROVER architectures
  saved a couple of `act_func_*` weights that modern KERMTEmbedding doesn't
  use. Benign; the rest of the encoder loaded correctly.

## What's in `run.json` after a successful run

Same reproducibility fields as `kermt-continue-pretrain`, plus the upgrade step's
`summary.json` is captured under the `inputs.upgrade_summary` path so the
provenance of the upgraded ckpt is auditable.

## Replayability

Same as `kermt-continue-pretrain`: `cmd_replay` rebuilds the
`run_pretrain_local.py --ckpt <upgraded.pt> ...` invocation. To redo the
full add-cmim flow end-to-end, the user also needs the input grover_base
ckpt and the corpus — both are captured in the prepare_data and upgrade
manifests by absolute path.
kermt-continue-pretrain15.1 KB

View saved version →

---
name: kermt-continue-pretrain
description: Continue pretraining from an existing KERMT checkpoint. The skill validates the user's checkpoint and pretrain CSV, prepares the data into shard/vocab/features form, then launches pretrain_ddp.py inside the kermt container (detached for long runs). Auto-dispatches `--pretrain_mode` based on the checkpoint type (grover_base vocab-only, cmim, or hybrid).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: this file is targeted at ~250 lines / ~3000 tokens —
# well within the 500-line / 5000-token cap. Long examples live in
# agent/scripts/run_pretrain_local.py's docstring.
---

# kermt-continue-pretrain

Continue pretraining from a user-supplied KERMT checkpoint (grover_base /
cmim / hybrid). The skill is the workflow orchestrator: it validates inputs,
prepares the corpus, launches the runner, and returns a run directory.

## Hardware requirements

- **GPUs**: 1–N CUDA-capable NVIDIA GPUs. The runner auto-detects via
  `torch.cuda.device_count()`; `--gpus 0,2` overrides. On a single GPU the
  runner falls back to `--batch_size 32 --save_interval 500`; on multi-GPU
  it uses the `defaults_pretrain.json` values (currently `batch_size 256`).
  Note: `--gpus N` uses **torch.cuda** indexing, which can differ from
  `nvidia-smi`'s display order on multi-GPU hosts (PCI bus vs. CUDA
  enumeration). To target a specific physical GPU, set `CUDA_VISIBLE_DEVICES`
  before invoking, or run
  `python -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])"`
  to confirm which device you're picking.
- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware
  (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:

  | GPU class                | VRAM       | Suggested `--batch-size` |
  |--------------------------|------------|--------------------------|
  | L4, T4, V100 16 GB       | 16–24 GB   | 32–64                    |
  | A100 40 GB, L40, A40     | 40–48 GB   | 128                      |
  | A100 80 GB, H100, H200   | 80 GB      | 256 (default)            |

  These are rough starting points — pass `--batch-size N` to override.
- **Disk**: tens of GB depending on corpus size + epochs (each checkpoint
  is several hundred MB).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).
  `kermt-setup` validates this up-front.

## Inputs

Required:

- `--csv <path>` — the pretrain CSV (single column `smiles`). If you have
  separate train/val CSVs, pass `--val-csv <path>` too.

Checkpoint (optional — defaults to the released model if omitted):

- `--ckpt <path>` — the input pretrain checkpoint to continue from. Must be
  a grover_base (with vocab heads), cmim, or hybrid ckpt; the validator
  rejects everything else with a redirect to the correct workflow. **If
  omitted**, the skill offers to download the released pretrained hybrid model
  **nvidia/NV-KERMT-70M-v2** and continue-pretrain from it — see "Resolve &
  validate the checkpoint" (workflow step 3). The released bundle ships its
  three vocab files alongside the ckpt, so the authoritative-vocab pass-through
  (step 5) works automatically.
- `--pretrained-release` — explicit opt-in to use the released model without
  the interactive prompt (for non-interactive / agent runs). Mutually
  exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
  reused, not re-downloaded.

Optional:

- `--val-csv <path>` — separate validation CSV. Without it, the prep step
  auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).
- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /
  `--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /
  `--save-interval N` / `--seed N` — training-hyperparameter overrides.
  Anything not given is filled from `agent/config/defaults_pretrain.json`.
- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /
  `--contrastive-temperature F` (cmim and hybrid only) — loss / decoder
  overrides.
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
  logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run
  name is honored only alongside a project. Off by default. (Independent of the
  ckpt's `wandb_run_id` continuity handling under `--resume`.)
- `--resume` — see "Modes" section below.
- `--gpus 0,2` — restrict to a GPU subset. Default uses all visible GPUs.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
  `prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.

## Modes

The runner has two modes for ingesting the input ckpt, dispatched on whether
`--resume` is set. Pick based on intent:

### Default (fresh-schedule continue-pretrain)

**Use when**: you have a finished pretrain ckpt and want to continue training
it — on a new corpus, with a different objective, or just for more epochs
than its original plan. The previous training's step counter and schedule
shape are no longer relevant; you want a new learning-rate schedule for the
new run.

**What gets loaded from the ckpt**:
- ✓ Model weights (encoder + vocab heads + contrast head + decoder, whatever
  is there)
- ✓ Optimizer state (Adam's running m1/m2 moments — warm-starts the new
  schedule so the first few hundred steps aren't dominated by noisy
  gradient-estimate startup)
- ✗ Scheduler step counter (reset to 0)
- ✗ Epoch counter (reset to 0)
- ✗ Batch counter (reset to 0)
- ✗ wandb run id (new wandb run, not a continuation)

**Schedule shape** (init/max/final LR, warmup epochs, total epochs): from
your CLI args or `defaults_pretrain.json`. A fresh NoamLR is constructed
from these values and starts at step 0.

### `--resume` (true resume)

**Use when**: a previous run was interrupted (crash, OOM, Ctrl-C) and you
want to pick up exactly where it left off — same dataset, same schedule,
same training trajectory.

**What gets loaded from the ckpt**: **everything** in the
`save_model_for_restart` format. Model weights + optimizer state +
scheduler_step + epoch + batch_idx + wandb_run_id are all restored. The
new run continues from the saved step in the saved schedule (which is
recovered from the ckpt's `saved_args`). Mid-epoch resume works too —
`pretrain_ddp.py`'s sampler skip-count picks up at the saved batch index
within the saved epoch.

**Schedule shape**: inherited from the ckpt's `saved_args`. CLI overrides
of any schedule flag (`--epochs / --warmup-epochs / --init-lr / --max-lr /
--final-lr`) are **rejected with a hard error** — pure resume means pure
resume; if you want to change the schedule, drop `--resume` and start a
fresh-schedule run.

**Requirements**: the ckpt must have been saved via `save_model_for_restart`
(i.e., carry `optimizer / scheduler_step / epoch / batch_idx` keys). If
any of these is missing, the runner errors with a clear message and
suggests dropping `--resume`.

The default mode is the right choice ~90% of the time. Reach for `--resume`
only when you genuinely need to continue a single interrupted training
run.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has already built `kermt:latest`. All paths below are on the
host; the helper bind-mounts them at known container paths.

1. **Pre-flight: ensure container + system probe.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c "
   import json, sys; d = json.load(sys.stdin)
   if not d['ok']:
       print('System check failed:', d['gaps']); sys.exit(1)
   print(f'OK: {len(d[\"gpus\"])} GPU(s); {d[\"disk\"][\"free_gb\"]} GB free; CUDA via container toolkit')
   "
   ```
   Surface any `gaps` to the user. Refuse to proceed if `ok: false`.

2. **Compute run directory.**
   ```
   RUN_DIR=$KERMT_REPO/runs/continue-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Resolve & validate the checkpoint.**

   **Resolve — only if `--ckpt` was omitted.** Default to the released
   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
     "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
     and continue-pretrain from it? [y/N]". **Never download without an
     explicit yes** (or `--pretrained-release`). If both `--ckpt` and
     `--pretrained-release` are given, abort — they conflict.
   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
     `--model-dir <dir>` if given. An already-complete bundle is reused.
   - **Download** (foreground; ~282 MB on first fetch):
     ```
     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
         "python agent/scripts/fetch_released_model.py --out /model"
     ```
     Parse the JSON; abort on `ok: false` (surface `errors`). On success set
     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`. The bundle's three
     vocab files land in `<save-dir>` too, so step 5's `--vocab-dir`
     auto-detection (which looks in the ckpt's parent directory) finds them
     with no extra work.

   **Validate** the resolved (or user-provided) ckpt:
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
       "python agent/scripts/check_checkpoint.py --mode continue_pretrain --ckpt /ckpt"
   ```
   Parse the JSON. Abort on `ok: false`, showing the error verbatim. The error
   message redirects the user to `kermt-add-cmim-pretrain` for encoder-only
   ckpts, or to `kermt-finetune` for finetuned ckpts.

4. **Validate the data.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
       "python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>"
   ```
   Abort on `ok: false`.

5. **Prepare the data** (skip if `--from-prepare` given).
   **Pass the ckpt's vocab through.** Look in the ckpt's parent directory for
   the conventional `pretrain_atom_vocab.{json,pkl}`, `pretrain_bond_vocab.{json,pkl}`,
   and `pretrain_smiles_vocab.pkl` files (the bundling convention for released
   models; see `agent/README.md` "Released models" section). If all three are
   present, auto-pass via `--vocab-dir <ckpt_parent_dir>`. If only some are
   present, pass them via explicit flags (`--atom-vocab`, `--bond-vocab`,
   `--smiles-vocab`). If none are present, ask the user for `--vocab-dir` — or
   refuse to proceed, because rebuilding a fresh vocab from the new corpus
   would silently mismatch the ckpt's vocab heads (the ckpt's vocab is
   authoritative for continue-pretrain).

   Note the **two-layer mount pattern**: pass the host directory to
   `kermt_container.sh --vocab-dir` (which mounts it at `/vocab` inside the
   container), and reference `/vocab` from the inner `prepare_data.py`
   command. The same pattern applies to every host path the inner command
   needs to read (`--data <host-csv>` → `/data/<basename>`,
   `--ckpt <host-ckpt>` → `/ckpt`).

   ```
   VOCAB_DIR=$(dirname <user-ckpt>)
   $KERMT_REPO/agent/scripts/kermt_container.sh run \
       --data <user-csv> --vocab-dir $VOCAB_DIR --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode pretrain \\
            --csv /data/<basename> --out /runs/data \\
            --vocab-dir /vocab \\
            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
   ```
   Outputs land at `$RUN_DIR/data/prepare_data.json` with
   `vocab_source: "user_provided"`. The runner step 7 will verify the vocab
   files' entry counts match the ckpt's vocab-head sizes and refuse to launch
   on mismatch.

6. **Estimate runtime + confirm with user.**
   - Pretrain wall time depends on corpus size × epochs × GPU count.
   - Tell the user the estimate; ask "proceed?" unless `--yes` flag was given
     (agent-non-interactive case).
   - Example estimate template:
     `~N hours on K GPUs for E epochs over M molecules (~steps/epoch × seconds/step)`.

7. **Launch the runner detached.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
       --name kermt-continue-pretrain-<ts> \\
       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_pretrain_local.py \\
            --ckpt /ckpt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs \\
            [--epochs N --batch-size N --init-lr F ...]"
   ```
   Returns the container name + id + log file path.

8. **Report to the user.** Output a short summary:
   - Container name + id
   - `$RUN_DIR/run.json` (the manifest with cmd_replay + image digest)
   - Log file: `$RUN_DIR/logs/pretrain_ddp.log`
   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
     $RUN_DIR/logs/tb`)
   - Suggest invoking `kermt-monitor <RUN_DIR>` to check progress.

## Hard rules

- **Never download the released model without consent.** When `--ckpt` is
  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
  or an explicit `--pretrained-release` flag. `--ckpt` and
  `--pretrained-release` are mutually exclusive.
- **Never modify the user's input ckpt.** The runner symlinks it into the
  save_dir; the symlink is what pretrain_ddp.py auto-resumes from. The
  source file stays untouched.
- **Never silently override arch.** If the user passes a `--hidden-size`
  etc. that doesn't match the ckpt-derived value, the runner aborts loudly.
  Arch params come from the ckpt, period.
- **Never block on the long-running pretrain itself.** The runner is invoked
  via `run_detached`; the skill returns immediately after step 8. Use
  `kermt-monitor` for progress.
- **Echo applied defaults back to the user.** The `args_applied` field of
  `run.json` records every flag's value + source (user / default-config /
  auto-1gpu / auto-multi-gpu). Skill should surface a summary of any flag
  not user-specified so the user knows what was assumed.

## Common errors

- `model_type='finetuned'` rejected → the ckpt is a downstream finetune,
  not a pretrain. The error redirects to the relevant workflow.
- `grover_base ckpt has no vocab head` → encoder-only ckpt (e.g. the
  original-grover `grover_base.pt`). The error redirects to
  `kermt-add-cmim-pretrain`.
- `prepare_data manifest is missing required outputs` → user passed
  `--from-prepare` to a directory where prepare was run with `--skip-vocab`
  or `--skip-split`. Re-run prepare without those flags.
- `--gpus all` not available → install `nvidia-container-toolkit`; check
  `kermt_container.sh check_system`.

## Replayability

The `run.json` `cmd_replay` field is a single-line command that re-runs the
pretrain with the same inputs, hyperparameters, and arch. To replay:

```bash
# Inside the kermt container:
$(jq -r .cmd_replay $RUN_DIR/run.json)
```

If `ok_to_replay: false` in the manifest (because the kermt repo working
tree was dirty at launch time), the replay may not be bit-exact — pin the
exact commit via the `repo.commit` field and `git checkout` it
first.
kermt-embed6.6 KB

View saved version →

---
name: kermt-embed
description: Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (grover_base / cmim / hybrid / finetuned). Writes one .npy per readout type (atom_from_atom, bond_from_atom, atom_from_bond, bond_from_bond) plus canonical_smiles.npy and validity.npy. Calls task/extract_embeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: targets ~150 lines / ~1800 tokens — within the
# 500-line / 5000-token cap for skill files.
---

# kermt-embed

Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.
The skill is the workflow orchestrator: validate ckpt, validate CSV, clean
SMILES, launch the runner blocking, return the per-readout `.npy` files.

## Hardware requirements

- **GPUs**: 1 (single-GPU).
- **VRAM**: ≥ 4 GB for the default `batch_size 64`.
- **Disk**: depends on output size — roughly a few MB per 1k molecules at
  `hidden 800` per readout, so ~10–20 MB per 1k molecules across the 4
  readouts. Plus a small `canonical_smiles.npy` + `validity.npy` per run.
- **Driver / CUDA**: any host supporting CUDA 12.6.

## Inputs

Required:

- `--csv <path>` — SMILES CSV. First column is `smiles`; other columns
  are ignored (no targets needed).

Checkpoint (optional — defaults to the released model if omitted):

- `--ckpt <path>` — any encoder-bearing checkpoint. Grover_base, cmim,
  hybrid, and finetuned ckpts are all accepted. The validator only refuses
  ckpts with no encoder. **If omitted**, the skill offers to download the
  released pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and embed with
  it — see "Resolve & validate the checkpoint" (workflow step 3).
- `--pretrained-release` — explicit opt-in to use the released model without
  the interactive prompt (for non-interactive / agent runs). Mutually
  exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
  reused, not re-downloaded.

Optional:

- `--batch-size N` — override the configured default (64).
- `--gpus 0` — single GPU id (default 0).
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
  `prepare_data.json` in `<dir>`.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout.

1. **Pre-flight: container + system probe.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh check_system
   ```

2. **Compute run directory.**
   ```
   RUN_DIR=$KERMT_REPO/runs/embed_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Resolve & validate the checkpoint.**

   **Resolve — only if `--ckpt` was omitted.** Default to the released
   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
     "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
     and embed with it? [y/N]". **Never download without an explicit yes** (or
     `--pretrained-release`). If both `--ckpt` and `--pretrained-release` are
     given, abort — they conflict.
   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
     `--model-dir <dir>` if given. An already-complete bundle is reused.
   - **Download** (foreground; ~282 MB on first fetch):
     ```
     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
         "python agent/scripts/fetch_released_model.py --out /model"
     ```
     Parse the JSON; abort on `ok: false` (surface `errors`). On success set
     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.

   **Validate** the resolved (or user-provided) ckpt:
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
       "python agent/scripts/check_checkpoint.py --mode embed --ckpt /ckpt"
   ```
   Parse JSON. Abort on `ok: false`. The validator only refuses encoder-less
   ckpts (rare).

4. **Validate the data.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
       "python agent/scripts/check_data.py --mode embed --csv /data/<basename>"
   ```

5. **Prepare the data** (clean-only — no features step).
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode embed \\
            --csv /data/<basename> --out /runs/data"
   ```
   Outputs land at `$RUN_DIR/data/prepare_data.json` with a single `clean_csv`
   path. `task/extract_embeddings.py` featurizes from SMILES on the fly.

6. **Launch the runner (blocking).**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run \\
       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_extract_embeddings.py \\
            --ckpt /ckpt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs \\
            [--gpus 0 --batch-size N]"
   ```

7. **Report to the user.**
   - Embeddings directory: `$RUN_DIR/out/`
     - `atom_from_atom.npy`, `bond_from_atom.npy`,
       `atom_from_bond.npy`, `bond_from_bond.npy` (the 4 standard readouts;
       each shape `(N_rows, hidden_size)`)
     - `metadata.pkl` — pickle of a dict containing `canonical_smiles`
       (RDKit-canonicalized SMILES per row), `valid` (boolean per-row: did
       RDKit parse it), plus other run metadata.
   - Manifest: `$RUN_DIR/run.json`
   - Log: `$RUN_DIR/logs/embed.log`

## Hard rules

- **Never download the released model without consent.** When `--ckpt` is
  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
  or an explicit `--pretrained-release` flag. `--ckpt` and
  `--pretrained-release` are mutually exclusive.
- **Never modify the user's ckpt.** The runner reads-only via
  `task/extract_embeddings.py`'s `--checkpoint <path>` flag.
- **Arch comes from the ckpt.** No `--hidden-size` flag etc. on this runner;
  `task/extract_embeddings.py` reads arch from the ckpt's saved_args.

## Common errors

- `prepare_data manifest is missing required output 'clean_csv'` → prepare
  ran with `--skip-clean` but no source CSV given. Re-run prepare without it.
- `--gpus '0,1' is single-GPU only` → pass a single id.

## Replayability

```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```

If `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin
the commit via `repo.commit` and `git checkout` it first.
kermt-finetune15.2 KB

View saved version →

---
name: kermt-finetune
description: Finetune a pretrained KERMT encoder on a labeled CSV. The skill validates the input checkpoint (must be a pretrain ckpt — grover_base / cmim / hybrid), validates the labeled CSV, prepares the data (clean + features + optional split), then launches main.py finetune inside the kermt container (detached for hours-scale runs). Hyperparameters come from agent/config/defaults_finetune.json with per-flag CLI override.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: targets ~250 lines / ~3000 tokens — within the
# 500-line / 5000-token cap for skill files.
---

# kermt-finetune

Finetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill
is the workflow orchestrator: validate ckpt, validate data, prepare data,
launch the runner detached, return a run directory + container name.

## Hardware requirements

- **GPUs**: 1 by default (single-GPU); pass `--gpus 0` (or whichever id) to
  select one. For faster training on a multi-GPU host, pass `--num-gpus N`
  (N>1) to run data-parallel DDP across N GPUs — `--batch-size` is then
  per-GPU (effective global batch = batch_size × N).
- **VRAM**: ≥ 8 GB for the default `batch_size 32` configuration. Lower VRAM
  works at smaller batch sizes — pass `--batch-size N` to override.
- **Disk**: a few GB per run (checkpoint + features + logs).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).
  `kermt-setup` validates this up-front.

## Inputs

Required:

- `--csv <path>` — labeled CSV. First column is `smiles`; every other column
  is a target.

Checkpoint (optional — defaults to the released model if omitted):

- `--ckpt <path>` — input pretrain checkpoint (grover_base / cmim / hybrid).
  The validator refuses already-finetuned ckpts with a redirect to
  `kermt-infer`. **If omitted**, the skill offers to download the released
  pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and finetune from it —
  see "Resolve & validate the checkpoint" (workflow step 3).
- `--pretrained-release` — explicit opt-in to use the released model without
  the interactive prompt (for non-interactive / agent runs). Mutually
  exclusive with `--ckpt`.
- `--model-dir <dir>` — where to save the downloaded bundle (default
  `$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is
  reused, not re-downloaded.

Optional:

- `--dataset-type {regression | classification | multiclass}` — default
  `regression` (from `defaults_finetune.json`). Drives loss, metric defaults,
  and head initialization. For classification tasks pass
  `--dataset-type classification`.

- `--targets COL [COL ...]` — explicit target column names. If omitted, the
  validator auto-detects numeric non-smiles columns and the skill confirms
  with the user before proceeding.
- `--val-csv <path>` and `--test-csv <path>` — user-provided val + test
  splits. Either pass both or pass neither (the skill auto-splits using the
  configured `--split-type`).
- `--split-type {random | scaffold_balanced | index_predetermined}` —
  default `scaffold_balanced` from `defaults_finetune.json`.
  - `random` and `scaffold_balanced`: build the val/test split internally
    from the train CSV. No `--val-csv` / `--test-csv` needed.
  - `index_predetermined`: **requires** pre-split CSVs passed via
    `--val-csv` + `--test-csv` (and, separately, per-fold index files —
    see `kermt/util/utils.split_data`). Use this when the dataset ships
    its own canonical split (e.g. `tests/data/Biogen_for_grover/scaffold/
    balance/<endpoint>/{train,val,test}.csv`).
- `--metric NAME` — `mae` (regression default), `auc` (classification default),
  or any name `kermt.util.metrics.get_metric_func` accepts.
- `--epochs N` / `--batch-size N` / `--init-lr F` / `--max-lr F` /
  `--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` /
  `--bond-drop-rate F` / `--dist-coff F` / `--early-stop-epoch N` /
  `--seed N` — training-hyperparameter overrides. Anything not given is
  filled from `agent/config/defaults_finetune.json`.
- `--ffn-hidden-size N` / `--ffn-num-layers N` — shared FFN trunk dims.
- `--ffn-num-task-specific-layers N` / `--ffn-task-specific-hidden-size H` —
  per-target FFN heads (default 0 = off; useful for heterogeneous multi-target
  finetunes). Both must be set together when N > 0.
- `--ensemble-size N` / `--num-folds N` — multi-model / k-fold CV. Default 1
  each.
- `--gpus 0` — single GPU id for single-process finetune (default 0). Ignored
  when `--num-gpus > 1`.
- `--num-gpus N` — number of GPUs for data-parallel DDP finetune. Default 1
  (single-process, unchanged). N>1 runs `main.py finetune` with `WORLD_SIZE=N`
  (one process per GPU); `--batch-size` is per-GPU.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
  `prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has built `kermt:latest`. All paths below are on the host; the
helper bind-mounts them at known container paths.

1. **Pre-flight: ensure container + system probe.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh check_system | python -c "
   import json, sys; d = json.load(sys.stdin)
   if not d['ok']:
       print('System check failed:', d['gaps']); sys.exit(1)
   print(f'OK: {len(d[\"gpus\"])} GPU(s); CUDA via container toolkit')
   "
   ```
   Refuse to proceed if `ok: false`.

2. **Compute run directory.**
   ```
   RUN_DIR=$KERMT_REPO/runs/finetune_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Resolve & validate the checkpoint.**

   **Resolve — only if `--ckpt` was omitted.** Default to the released
   pretrained hybrid model **nvidia/NV-KERMT-70M-v2**:
   - **Consent gate.** Unless `--pretrained-release` was passed, ask the user:
     "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2
     (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2)
     and finetune from it? [y/N]". **Never download without an explicit yes**
     (or `--pretrained-release`). If both `--ckpt` and `--pretrained-release`
     are given, abort — they conflict.
   - **Save location.** Default `$KERMT_REPO/models/NV-KERMT-70M-v2/`; honor
     `--model-dir <dir>` if given. An already-complete bundle is reused.
   - **Download** (foreground; ~282 MB on first fetch):
     ```
     $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
         "python agent/scripts/fetch_released_model.py --out /model"
     ```
     Parse the JSON; abort on `ok: false` (surface `errors`). On success set
     `<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt`.

   **Validate** the resolved (or user-provided) ckpt:
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
       "python agent/scripts/check_checkpoint.py --mode finetune_init --ckpt /ckpt"
   ```
   Parse the JSON. Abort on `ok: false`. The validator rejects already-
   finetuned ckpts (`has_task_ffn: true`) with a redirect to `kermt-infer`.

4. **Validate the data.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
       "python agent/scripts/check_data.py --mode finetune --csv /data/<basename> [--targets COL1 COL2 ...]"
   ```
   If `--targets` was not given by the user, surface `auto_detected_targets`
   from the JSON and ask the user to confirm before continuing. Abort on
   `ok: false`.

5. **Prepare the data** (skip if `--from-prepare` given).

   **Pre-flight: check for sibling val.csv / test.csv.** Before invoking
   prepare_data, inspect the parent directory of `<user-csv>`. If a
   canonical-looking sibling `val.csv` (or `val_*.csv` — common variants
   include `val_T.csv`, `val_clean.csv`) AND a matching `test.csv` /
   `test_*.csv` exist next to the train CSV, the dataset ships its own
   pre-defined split. **In that case set `--split-type index_predetermined`
   AND pass `--val-csv` / `--test-csv`** — otherwise the configured
   `split_type` (default `scaffold_balanced`) will re-split the train CSV
   from scratch and silently discard the user's val/test files. When in
   doubt — or when the sibling files use non-canonical suffixes (`_T`,
   `_v2`, etc.) — surface the situation to the user and ask which they
   want.

   **Quoting target names.** If any of the `--targets` column names
   contain shell metacharacters (`>`, `&`, `|`, `(`, `)`, `$`, etc.),
   single-quote each one when passing on the CLI to keep the shell from
   eating part of the name. Example: `--targets 'Log_Caco2_Papp_A>B'
   'logD'`. The CSV header itself is read directly by the downstream
   trainer and is unaffected, but the prepare_data.json manifest's
   `targets[]` field captures whatever the shell delivers — unquoted
   metacharacters get truncated there.

   **Mount note:** `kermt_container.sh --data <host-csv>` mounts the
   parent directory of `<host-csv>` at `/data`. `--val-csv` and
   `--test-csv` must therefore reference files in that same parent
   directory. If val/test live in a separate directory (e.g. a sibling
   `splits/` folder), mount the parent of all three using `--data <dir>`
   on a directory rather than a file.

   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode finetune \\
            --csv /data/<basename> --out /runs/data \\
            --split-type <split_type> \\
            [--val-csv /data/<val-basename> --test-csv /data/<test-basename>] \\
            [--val-frac 0.1 --test-frac 0.1 --seed 0] \\
            --targets <COL1> [COL2 ...]"
   ```
   Outputs land at `$RUN_DIR/data/prepare_data.json`. For `scaffold_balanced`
   and `index_predetermined`, prep emits a single `clean_full_csv` + `.npz`;
   the runner passes them through to `main.py finetune` which calls
   `split_data` internally with the user-supplied seed.

6. **Estimate runtime + echo applied defaults.**
   - Finetune wall time is typically minutes-to-hours on 1 GPU.
   - Surface a summary of every flag that was filled from the defaults
     vs user-supplied, so the user knows what was assumed. The runner
     records this in `args_applied`.
   - Sample message:
     `"Filling from defaults_finetune.json: epochs=30, batch_size=32,
       split_type=scaffold_balanced. Override any of these with --<flag>."`

7. **Targets confirmation gate (hard requirement).** Before launching the
   runner, regardless of how the targets list was determined (CLI `--targets`,
   auto-detection in step 4, or a user natural-language request like
   "finetune on Caco2 and HLM"), echo the final targets list to the user with
   an explicit count:
   `"Will finetune on N target(s): COL1, COL2, ..."`. If the user's request
   specified a subset that doesn't match this list (e.g., they asked for 2
   tasks via natural language but the list still has 4), treat it as a
   discrepancy and re-prompt with the diff — never silently proceed on the
   wrong target set. Wait for explicit confirmation before launching unless
   `--yes` was given.

8. **Launch the runner detached.** (Consistent with the pretrain skills.)
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
       --name kermt-finetune-<ts> \\
       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_finetune_local.py \\
            --ckpt /ckpt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --dataset-type <type> \\
            --out /runs \\
            [--gpus 0] \\
            [--num-gpus N] \\
            [--epochs N --batch-size N --init-lr F ...] \\
            [--ffn-num-task-specific-layers N --ffn-task-specific-hidden-size H]"
   ```
   Returns the container name + id + log file path.

9. **Report to the user.** Output a short summary:
   - Container name + id
   - `$RUN_DIR/run.json` (manifest with cmd_replay + image digest)
   - Log file: `$RUN_DIR/logs/finetune.log`
   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
     $RUN_DIR/logs/tb`)
   - Final checkpoints land at `$RUN_DIR/ckpt/fold_0/model_0/model.pt`
     (best-val) and `last_checkpoint.pt` (sibling, auto-resume target).
     Held-out test predictions + metrics land at
     `$RUN_DIR/ckpt/fold_0/test_result.csv`. Paths vary with `--num-folds`
     / `--ensemble-size`.
   - To follow progress: `kermt-monitor <RUN_DIR>` (one-shot) or
     `docker logs -f <container-name>` (streaming).
   - To block until the run finishes (useful for short test runs):
     `docker wait <container-name>` — prints the exit code on completion.

## Hard rules

- **Never download the released model without consent.** When `--ckpt` is
  omitted, download `nvidia/NV-KERMT-70M-v2` only after an explicit user "yes"
  or an explicit `--pretrained-release` flag. `--ckpt` and
  `--pretrained-release` are mutually exclusive.
- **Never modify the user's input ckpt.** The runner passes its path via
  `--checkpoint_path`; `task/train.py` loads it read-only into the model and
  attaches a new FFN head. The source file stays untouched.
- **Arch comes from the ckpt, not from CLI/defaults.** The runner extracts
  `hidden_size`, `depth`, `num_attn_head`, `activation`, `embedding_output_type`,
  `self_attention` (+ `attn_hidden` / `attn_out` when applicable) from the
  ckpt's saved_args. There is no `--hidden-size` flag on this runner.
- **Never block on the long-running finetune.** The skill launches via
  `run_detached` and returns immediately after step 9. Use `kermt-monitor`.
- **Echo applied defaults back to the user.** The `args_applied` field of
  `run.json` records every flag's value + source (user / default-config).
  Surface a one-line summary of every filled-from-default flag so the user
  knows what was assumed.

## Common errors

- `finetune_init requires a pretrain ckpt (grover_base / cmim / hybrid)` →
  the ckpt you passed is already finetuned (has task FFN heads). Pick a
  pretrain ckpt instead, or use `kermt-infer` if you want to run
  predictions with the existing finetuned model. To resume a finetune on
  the SAME dataset, bypass the skill and call
  `python main.py finetune --checkpoint_path <ckpt> ...` directly — the
  agent skill doesn't support resume because saved-task identity
  can't be machine-verified against the new training data.
- `prepare_data manifest reports ok=False` → check `errors` for the failed
  step (typically clean_smiles or save_features). Fix and re-run.
- `ffn_num_task_specific_layers=N>0 but ffn_task_specific_hidden_size is unset`
  → MTL heads need an explicit hidden size. Pass `--ffn-task-specific-hidden-size H`.
- `finetune is single-GPU` (from `--gpus 0,1`) → `--gpus` selects one device
  for single-process finetune. For multi-GPU, use `--num-gpus N` (DDP) instead.

## Replayability

The `run.json` `cmd_replay` field is a single-line command that re-runs the
finetune with the same inputs, hyperparameters, and arch. To replay inside
the kermt container:

```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```

If `ok_to_replay: false` in the manifest (because the kermt repo working
tree was dirty at launch time), the replay may not be bit-exact — pin the
exact commit via the `repo.commit` field and `git checkout` it
first.
kermt-infer5.56 KB

View saved version →

---
name: kermt-infer
description: Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The skill validates that the input ckpt has task FFN heads (refuses pretrain ckpts with a redirect to kermt-finetune), validates the CSV, prepares the data (clean + rdkit_2d features), then launches main.py predict inside the kermt container (blocking, minutes-scale).
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: targets ~170 lines / ~2000 tokens — well within the
# 500-line / 5000-token cap for skill files.
---

# kermt-infer

Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. The
skill is the workflow orchestrator: validate ckpt, validate CSV, prepare data,
launch the runner blocking, return the predictions CSV.

## Hardware requirements

- **GPUs**: 1 (single-GPU). Multi-GPU inference is not currently supported.
- **VRAM**: ≥ 4 GB for the default `batch_size 32`.
- **Disk**: a few hundred MB per run (cleaned CSV + features + predictions).
- **Driver / CUDA**: any host supporting CUDA 12.6 (the kermt image base).

## Inputs

Required:

- `--ckpt <path>` — finetuned checkpoint (must have task FFN heads). The
  validator refuses pretrain ckpts with a redirect to `kermt-finetune`.
- `--csv <path>` — SMILES-only CSV. First column is `smiles`; other columns
  are ignored.

Optional:

- `--batch-size N` — override the configured default (32).
- `--seed N` — random seed for inference (deterministic featurization paths).
- `--gpus 0` — single GPU id (default 0). Multi-GPU rejected.
- `--from-prepare <dir>` — skip the prepare step and reuse an existing
  `prepare_data.json` in `<dir>`.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout, and assume
`kermt-setup` has built `kermt:latest`.

1. **Pre-flight: ensure container + system probe.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh check_system
   ```
   Refuse to proceed on `ok: false`.

2. **Compute run directory.**
   ```
   RUN_DIR=$KERMT_REPO/runs/infer_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Validate the checkpoint.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
       "python agent/scripts/check_checkpoint.py --mode inference --ckpt /ckpt"
   ```
   Parse the JSON. Abort on `ok: false`. The validator rejects pretrain ckpts
   (`has_task_ffn: false`) with a redirect to `kermt-finetune`.

4. **Validate the data.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
       "python agent/scripts/check_data.py --mode inference --csv /data/<basename>"
   ```
   Abort on `ok: false`.

5. **Prepare the data.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode inference \\
            --csv /data/<basename> --out /runs/data"
   ```
   Outputs land at `$RUN_DIR/data/prepare_data.json` with `clean_csv` +
   `clean_npz` paths (rdkit_2d_normalized features).

6. **Launch the runner (blocking).**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run \\
       --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_inference.py \\
            --ckpt /ckpt \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs \\
            [--gpus 0 --batch-size N --seed N]"
   ```
   Returns the predictions CSV path on success.

7. **Report to the user.** Output a short summary:
   - Predictions: `$RUN_DIR/out/predictions.csv` (smiles + per-target columns)
   - Manifest: `$RUN_DIR/run.json` (cmd_replay + image digest + applied args)
   - Log: `$RUN_DIR/logs/inference.log`
   - Row count: <N> molecules predicted across <K> targets

## Hard rules

- **Never modify the user's ckpt.** The runner symlinks the ckpt into a
  unique `<out>/ckpt_link/` subdir so `main.py predict --checkpoint_dir`
  picks it up; the source file stays untouched.
- **Arch comes from the ckpt, never from CLI/defaults.** The runner records
  the validator's arch block in `run.json` but does not pass arch flags into
  `main.py predict` — predict reads them from the loaded ckpt's saved_args.
- **Single-GPU only.** Multi-GPU inference is not currently supported.
- **Echo applied defaults.** The `args_applied` field of `run.json` records
  every flag's value + source (user / default-config). Surface a short
  summary of any default-filled flag.

## Common errors

- `inference requires a finetuned ckpt with task FFN heads` → ckpt is a
  pretrain ckpt; use `kermt-finetune` first.
- `prepare_data manifest reports ok=False` → check the manifest `errors` for
  the failed step (typically clean_smiles or save_features).
- `could not convert string to float: '<value>'` from save_features or main.py
  predict → input CSV has a non-numeric passthrough column (e.g. a 'split'
  label). The prep step now strips the CSV to SMILES-only at inference; if
  this error still surfaces, the CSV is being read by a runner that bypassed
  prepare_data. Re-run via the skill, not `main.py` directly.
- `--gpus '0,1' is single-GPU only` → pass a single id.

## Replayability

The `run.json` `cmd_replay` field is a single-line command that re-runs the
inference with the same inputs. To replay inside the kermt container:

```bash
$(jq -r .cmd_replay $RUN_DIR/run.json)
```

If `ok_to_replay: false` (dirty kermt repo worktree at launch time), pin
the commit via `repo.commit` and `git checkout` it first.
kermt-monitor6.95 KB

View saved version →

---
name: kermt-monitor
description: Check progress for a detached KERMT run (pretrain, finetune, or any kermt_run_detached invocation). Reads run.json, queries docker for container state, tails the pretrain/finetune log, and parses progress lines (epoch, step, val loss).
license: Apache-2.0
compatibility: Requires docker and jq. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: atomic-skill
  risk_tier: skill
# Line/token budget: this file is targeted at ~120 lines / ~1500 tokens — well
# within the 500-line / 5000-token cap for skill files.
---

# kermt-monitor

Companion skill for any KERMT workflow that runs detached: the three pretrain
skills (`kermt-continue-pretrain`, `kermt-pretrain-scratch`,
`kermt-add-cmim-pretrain`) plus `kermt-finetune`. `kermt-infer` and
`kermt-embed` run blocking by default and don't need this skill, but if a
user launches them detached on purpose the monitor still works (the
workflow-dispatch in step 4 handles unknown workflows by tailing the
most-recent log file in the run dir). Reads the run directory's `run.json`,
queries docker for the container's state, surfaces the latest progress,
and either tails or follows the log.

## Hardware requirements

None. This skill only reads disk + queries docker; no GPU compute.

## Inputs

One of:

- `<run-dir>` — a positional argument pointing at the directory containing
  `run.json` (e.g. `runs/continue-pretrain_2026-05-17T10-23Z`). Preferred.
- `--container <name-or-id>` — direct container reference; the skill still
  reads `run.json` from the run dir referenced inside the container's
  inspect output if available, but works degraded-mode without it.

Optional:

- `--lines N` — number of trailing log lines to print (default 50).
- `--follow` — stream `docker logs -f` until ^C. Useful for "watch the
  loss". Without it, the skill is one-shot and exits.
- `--json` — emit a structured status report instead of human-readable text.
  Useful when the parent agent wants to take downstream action.

## Workflow

Let `RUN_DIR=$1` (or whatever path the user supplies).

1. **Locate the manifest.**
   ```
   MANIFEST=$RUN_DIR/run.json
   ```
   Refuse to proceed if it doesn't exist; surface a helpful message
   pointing the user at the run-dir convention (`runs/<workflow>_<ts>/`).

2. **Parse the manifest** (Python helper):
   ```
   workflow=$(jq -r .workflow $MANIFEST)
   container_name=...   # not directly in run.json today; the skill that
                        # launched stored it in run.json under
                        # container.name during launch (see below note).
   logs_dir=$(jq -r .logs_dir $MANIFEST)
   image_tag=$(jq -r .container.image_tag $MANIFEST)
   started_at=$(jq -r .started_at $MANIFEST)
   ```

3. **Query docker for container state.**
   ```
   docker ps --filter "name=$container_name" --format \
       '{{.ID}}\t{{.Status}}\t{{.CreatedAt}}'
   ```
   If absent, fall back to `docker inspect $container_name --format
   '{{.State.Status}} (exit {{.State.ExitCode}})'` to see whether the
   container exited (ok or failed) or was removed (`--rm` after exit).

4. **Find the live log file.**
   ```
   case "$workflow" in
     continue-pretrain|pretrain-scratch)  LOG=$logs_dir/pretrain_ddp.log ;;
     finetune)                            LOG=$logs_dir/finetune.log ;;
     *)                                   LOG=$(ls -1t $logs_dir/*.log 2>/dev/null | head -n 1) ;;
   esac
   ```
   The manifest's `workflow` field disambiguates pretrain (`pretrain_ddp.log`)
   from finetune (`finetune.log`). Other workflows fall back to the
   most-recently-modified `.log` in `$logs_dir`.

5. **Show the latest progress.**
   - `tail -n $LINES $LOG` for the raw recent output.
   - Parse the last few progress lines and surface a human-friendly
     summary. The format differs per workflow:
     - Pretrain: epoch / step / val_loss
       ```
       Current epoch: 12/100  step: 4523/9000  val_loss: 0.832 (best 0.821 @ step 4100)
       ```
     - Finetune: fold / epoch / val_<metric> (e.g. val_mae for regression,
       val_auc for classification — read `args_applied.metric` from run.json)
       ```
       Fold 0  epoch 12/30  val_mae 0.187 (best 0.182 @ epoch 9)
       ```
     ```
     Wall-clock: 1h 23m since started_at; ETA ~6h remaining.
     ```

6. **Final test-metrics block (finetune, on completion).** If `workflow` is
   `finetune` AND the container has exited cleanly (`State.Status=exited`,
   `ExitCode=0`) AND `$RUN_DIR/ckpt/fold_*/test_result.csv` exists, parse it
   and emit a per-task metric table:
   ```
   Final test metrics (per task):
     Target              MAE
     HLM_clearance       0.187
     RLM_clearance       0.213
     MDR1-MDCK_efflux    0.241
     solubility_pH6.8    0.156
   ```
   The metric column matches `args_applied.metric` (mae for regression, auc
   for classification, etc.). For multi-fold or ensemble runs, average across
   folds/models and note `± std` if std > 0. Skip silently if no
   `test_result.csv` exists (run incomplete or no test split was emitted).

7. **If `--follow`, stream live logs.**
   ```
   docker logs -f $container_name
   ```
   Wraps until ^C.

8. **Stop / cleanup hints** (printed at end of one-shot mode):
   ```
   To stop:        docker stop $container_name
   To remove:     docker rm $container_name
   To re-run:    `$(jq -r .cmd_replay $MANIFEST)`
   ```

## Hard rules

- **Read-only on the user's data.** Never modify `run.json`, never touch the
  container's checkpoint dir. The monitor only inspects.
- **Don't kill the container without explicit user instruction.** If the
  user asks to stop, run `docker stop`; if they ask to abandon, leave it
  running and just exit.
- **Don't pull or modify the kermt image.** The monitor only reads.
- **JSON output mode is non-interactive.** Skip the "press ^C to exit"
  prompts and emit a single JSON document so the parent agent can pipe it.

## Note on container_name plumbing

The run.json schema as currently written does not yet include the launched
container name — `kermt_run_detached` prints it to stdout but the runner
script doesn't capture it into run.json. The monitor falls back to a
filesystem-based lookup: list `runs/<workflow>_*/` directories and match by
mtime; or accept `--container <name>` explicitly. Follow-up: have the
launching skill record container name into run.json before exiting.

## Output (text mode, default)

```
KERMT continue-pretrain · runs/continue-pretrain_2026-05-17T10-23Z
  Container : kermt-continue-pretrain-…  (Up 1 hour, status: running)
  Image     : kermt:latest@sha256:…
  Repo      : 2fe00f9 (clean)
  Started   : 2026-05-17T10:23:14Z (1h 23m ago)
  Workflow  : continue-pretrain, pretrain_mode=hybrid, world_size=2

  Latest log (last 50 lines from $LOG):
    [Epoch 12/100] step 4523/9000 loss 0.832 lr 1.2e-4
    [val] step 4100 val_loss 0.821 (new best)
    ...

  Progress: epoch 12/100, ~12% done. ETA ~6h.
  TensorBoard: tensorboard --logdir $RUN_DIR/logs/tb
  Replay command: $(jq -r .cmd_replay $RUN_DIR/run.json)
```
kermt-pretrain-scratch8.99 KB

View saved version →

---
name: kermt-pretrain-scratch
description: Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: workflow-skill
  risk_tier: skill
# Line/token budget: ~210 lines, ~2400 tokens — well within the
# 500-line / 5000-token cap for skill files. Most of the orchestration is
# shared with kermt-continue-pretrain; the differences are documented below.
---

# kermt-pretrain-scratch

Pretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful
when you want to retrain a model on a custom chemistry domain rather than
extending one of the released checkpoints. **Significantly more expensive than
`kermt-continue-pretrain`** — no warm start, so the loss curves need to descend
from scratch over many epochs.

## Hardware requirements

Same as `kermt-continue-pretrain`:

- **GPUs**: 1–N CUDA-capable. The runner auto-detects via
  `torch.cuda.device_count()`; `--gpus 0,2` overrides. Single-GPU fallback:
  `--batch_size 32 --save_interval 500`. Multi-GPU keeps defaults
  (`--batch_size 256` etc.). Note: `--gpus N` uses **torch.cuda** indexing,
  which can differ from `nvidia-smi`'s display order on multi-GPU hosts
  (PCI bus vs. CUDA enumeration). To target a specific physical GPU, set
  `CUDA_VISIBLE_DEVICES` before invoking, or run
  `python -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])"`
  to confirm which device you're picking.
- **VRAM**: the default `--batch-size 256` is sized for A100-class hardware
  (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:

  | GPU class                | VRAM       | Suggested `--batch-size` |
  |--------------------------|------------|--------------------------|
  | L4, T4, V100 16 GB       | 16–24 GB   | 32–64                    |
  | A100 40 GB, L40, A40     | 40–48 GB   | 128                      |
  | A100 80 GB, H100, H200   | 80 GB      | 256 (default)            |

  These are rough starting points — pass `--batch-size N` to override.
- **Disk**: tens of GB for shards + vocab + checkpoints, scaled by epochs.
- **Wall time**: this is the big difference. Pretraining from scratch on an
  11M-mol corpus at 100 epochs typically takes **days even on a multi-GPU box**.
  The skill prints an estimate before launching; confirm with the user.

## When to invoke

- User wants to train a new model on a custom corpus (e.g. domain-specific
  chemistry that the released ckpts don't cover).
- User wants to reproduce a pretrain config end-to-end without depending on a
  released ckpt.

For continuing an existing released ckpt, use `kermt-continue-pretrain`. For
adding a cMIM decoder to an encoder-only grover_base ckpt, use
`kermt-add-cmim-pretrain`.

## Inputs

Required:

- `--csv <path>` — the pretrain corpus CSV with a `smiles` column. Single file
  by convention; multi-file corpora deferred. Use `--val-csv` for a separate
  validation set.
- `--pretrain-target-mode {vocab|cmim|hybrid}` — which pretrain objective to
  use. **No default** — must be set explicitly so the user makes an informed
  choice:
  - `vocab` — original GROVER-style atom + bond vocab prediction (encoder-only
    output, lightweight).
  - `cmim` — contrastive + SMILES reconstruction objective. Requires building
    a SMILES vocab from the corpus.
  - `hybrid` — both vocab and contrastive objectives jointly (the
    state-of-the-art config from the KERMT manuscript).

Optional:

- `--val-csv <path>` — separate validation CSV. Without it, prepare_data
  auto-splits the input by `--val-frac 0.1` (random shuffle with `--seed`).
- Training-hyperparameter overrides: `--epochs N` / `--batch-size N` /
  `--init-lr F` / `--max-lr F` / `--final-lr F` / `--warmup-epochs F` /
  `--weight-decay F` / `--dropout F` / `--save-interval N` / `--seed N`.
  Anything not given is filled from `agent/config/defaults_pretrain.json`.
- `--vocab-loss-weight F` (hybrid only) / `--latent-dim N` /
  `--contrastive-temperature F` (cmim and hybrid only).
- `--wandb-project NAME` / `--wandb-run-name NAME` — optional Weights & Biases
  logging. When `--wandb-project` is set, rank 0 logs train/val losses; the run
  name is honored only alongside a project. Off by default.
- `--gpus 0,2` — restrict to a GPU subset.

## Workflow

Let `$KERMT_REPO` be the path to your kermt repo checkout.

1. **Pre-flight: ensure container + system probe** (same as
   `kermt-continue-pretrain` step 1). Refuse to proceed if `check_system`
   reports gaps.

2. **Compute run directory.**
   ```
   RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ)
   ```

3. **Validate the corpus** (no ckpt to validate, so this is the only input
   check):
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
       "python agent/scripts/check_data.py --mode pretrain --csv /data/<basename>"
   ```
   Abort on `ok: false`.

4. **Prepare the data** — no vocab pass-through (we want fresh vocab from
   corpus):
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
       "python agent/scripts/prepare_data.py --mode pretrain \\
            --csv /data/<basename> --out /runs/data \\
            [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"
   ```
   Outputs land at `$RUN_DIR/data/prepare_data.json` with
   `vocab_source: "built_fresh"`.

5. **Estimate runtime + warn loudly.** This is critical for pretrain-from-scratch:
   - "Pretraining from scratch is days-scale even on multi-GPU; the released
     KERMT checkpoints were each trained on millions of molecules for hundreds
     of GPU-hours. If you mainly want to leverage existing knowledge for a
     downstream task, consider `kermt-continue-pretrain` from a released ckpt
     instead, which converges in hours instead of days."
   - Show the corpus size × epochs × GPU count → estimated wall time.
   - Ask for explicit confirmation unless `--yes` was given.

6. **Launch the runner detached.**
   ```
   $KERMT_REPO/agent/scripts/kermt_container.sh run_detached \\
       --name kermt-pretrain-scratch-<ts> \\
       --run-dir $RUN_DIR -- \\
       "python agent/scripts/run_pretrain_local.py \\
            --from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\
            --prepare-manifest /runs/data/prepare_data.json \\
            --out /runs \\
            [--epochs N --batch-size N ...]"
   ```
   Note: NO `--ckpt` flag (the runner refuses if both `--from-scratch` and
   `--ckpt` are given). The runner uses the `arch` group from
   `agent/config/defaults_pretrain.json` to size the model.

7. **Report to the user.** Always include all of the following — do not
   omit the TensorBoard line under output-length pressure:
   - Container name + id
   - `$RUN_DIR/run.json` (the manifest with `workflow: pretrain-scratch`,
     `from_scratch: true`, `vocab_check: null`, `arch` from defaults, full
     `cmd_replay`)
   - Log file: `$RUN_DIR/logs/pretrain_ddp.log`
   - TensorBoard: `$RUN_DIR/logs/tb` (open with `tensorboard --logdir
     $RUN_DIR/logs/tb`)
   - Suggest `kermt-monitor <RUN_DIR>` for progress.

## Hard rules

- **Never accept a `--ckpt` flag.** From-scratch is exclusive with input
  ckpt — the runner enforces this; the skill should too.
- **Never silently default `--pretrain-target-mode`.** This is a significant
  architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user
  if not given on the CLI.
- **Strong warning before launching.** From-scratch pretrain is the most
  expensive workflow. The user needs to know what they're committing to.

## Common errors

- `--pretrain-target-mode is required when --from-scratch is set` → user
  forgot the mode flag. Prompt.
- `--from-scratch is incompatible with --ckpt` → user provided both; ask which
  one they meant.
- `defaults_pretrain.json has no arch group` → repo state issue (should never
  happen on a fresh clone); points the user at running `kermt-setup` again.

## What's in the manifest after a from-scratch run

Same reproducibility fields as continue-pretrain (`repo.commit`, `kermt_image`,
`cmd_replay`, `args_applied`), plus:

- `workflow`: `"pretrain-scratch"`
- `from_scratch`: `true`
- `inputs.ckpt`: `null`
- `ckpt_symlink`: `null`
- `vocab_check`: `null` (not verified — vocab built from corpus is
  authoritative for from-scratch)
- `arch`: the values pulled from `agent/config/defaults_pretrain.json`'s
  `arch` group (with any future CLI overrides applied).

## Replayability

Same as continue-pretrain: `cmd_replay` is a copy-pasteable command. If
`ok_to_replay: false`, the kermt repo working tree was dirty at launch
time — check `repo.commit` and `git checkout` it first.
kermt-setup6.18 KB

View saved version →

---
name: kermt-setup
description: Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test inside the container. Every other kermt-* skill depends on this; invoke it first.
license: Apache-2.0
compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
metadata:
  owner: evax@nvidia.com
  classification: atomic-skill
  risk_tier: skill
# This file is intentionally short (~110 lines, ~1200 tokens) — well within the
# 500-line / 5000-token budget for skill files. Longer reference material lives
# alongside agent/scripts/kermt_container.sh.
---

# kermt-setup

Bootstrap the KERMT agent environment. Run this once on a fresh machine (or
after the Dockerfile or `environment.yml` changes) before invoking any other
`kermt-*` skill.

## Hardware requirements

- **GPU**: at least one CUDA-capable NVIDIA GPU visible to the host. The image
  is based on `nvidia/cuda:12.6.3-cudnn-devel-ubuntu22.04`, so the host driver
  must support CUDA 12.6. Verify with host `nvidia-smi` before invoking.
- **Host docker**: docker engine + nvidia-container-toolkit. Without the
  toolkit, `docker run --gpus all` will fail at step 2 of the workflow below.
- **Disk**: ≈ 50 GB free for the built kermt image (`docker image inspect
  --format '{{.Size}}'` reports ≈ 44 GB; the `docker images` Size column
  can show ~100 GB because it counts shareable buildx attestation layers
  that are deduplicated across images). Plan for ~50 GB of unique on-disk
  storage; add a comfortable buffer if you're also keeping build cache.
- **Memory**: the build itself peaks at ~4 GB RAM during conda env solve.
- This skill does not run training/inference workloads itself; per-workflow
  hardware requirements (VRAM, GPU count) are declared in the respective
  `kermt-<workflow>` skills.

## When to invoke

- User explicitly asks (`/kermt-setup`, "set up kermt", "build the kermt image",
  etc.).
- Or another `kermt-*` skill detected that the image does not exist and routed
  here. (Most other skills call `kermt_ensure_image` themselves, so this is
  usually only needed for the first-time setup, debugging, or a forced rebuild.)

## Inputs

The skill takes no required arguments. Optional overrides (via env vars before
invoking, or by setting them in the user's shell):

- `KERMT_IMAGE` — image tag to build/verify (default: `kermt:latest`).
- `KERMT_REPO` — host path of the kermt repo checkout (default: auto-derived
  from the script's location).

If the user has not specified a repo path and the current working directory is
not inside a kermt repo clone, ask for the repo path before proceeding.

## Workflow

All work goes through `agent/scripts/kermt_container.sh`. The script's
subcommand dispatch can be invoked directly without sourcing — that is the
preferred form for skill use.

Let `HELPER=$KERMT_REPO/agent/scripts/kermt_container.sh`.

1. **Verify docker is installed and the daemon is reachable.**
   ```
   $HELPER check_docker
   ```
   Exit 0 → continue. Non-zero → surface the error to the user (typically
   "docker not on PATH" or "daemon not reachable"); do not attempt step 2.

2. **Verify GPU passthrough works.**
   ```
   $HELPER check_gpu
   ```
   This runs `docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04
   nvidia-smi` and checks the exit status. Non-zero → tell the user to install
   `nvidia-container-toolkit` on the host and confirm a CUDA-capable NVIDIA GPU
   is visible to the host (`nvidia-smi` on the host should also work). Stop
   here; without GPU passthrough the kermt image will build but no workflow
   will run.

3. **Build or verify the kermt image.**
   ```
   $HELPER ensure_image
   ```
   If the image already exists, this returns immediately. Otherwise it builds
   from `$KERMT_REPO/Dockerfile`. **Warn the user before invoking** that the
   first build takes ~10–20 minutes on a typical workstation and streams build
   logs to the console. Do not run this in the background — the user wants to
   see progress and any build failures must surface immediately.

4. **GPU smoke test inside the container.** Quote the whole `python` command
   as a single string — the helper passes args through `bash -c "$*"`, so
   unquoted multi-word commands get re-parsed and any embedded quotes are
   collapsed.
   ```
   $HELPER run -- 'python -c "import torch; print(\"cuda_available:\", torch.cuda.is_available()); print(\"device_count:\", torch.cuda.device_count())"'
   ```
   Expected output: `cuda_available: True` and a positive `device_count`. If
   `cuda_available` is `False` despite step 2 passing, something is wrong with
   the container's CUDA wiring — report the full output to the user and stop;
   do not declare the environment ready.

5. **Summary to user.** Report:
   - Image tag and ID (`docker image inspect $KERMT_IMAGE --format '{{.Id}}'`).
   - Image size (`docker image inspect $KERMT_IMAGE --format '{{.Size}}'`).
   - GPU count detected inside the container.
   - "Ready" — the user can now invoke other `kermt-*` skills.

## Hard rules

- Do **not** pull or push docker images. The kermt image is built locally only.
- Do **not** auto-delete or prune older `kermt:*` tags without the user's
  explicit confirmation — the user may be running a finetune or pretrain in
  another container that depends on a specific tag.
- Do **not** modify the host's docker daemon configuration, daemon.json, or
  user-group membership.
- Do **not** modify the `Dockerfile` or `environment.yml` as part of this
  skill. If the build fails because of a Dockerfile issue, surface the error
  and stop; let the user decide whether to edit.
- Do **not** rebuild the image when it already exists (i.e. do not pass a
  `--no-cache` or `--pull` flag to ensure_image) unless the user explicitly
  asks for a forced rebuild.

## Forced rebuild

If the user explicitly asks to rebuild (e.g. after changing the Dockerfile or
`environment.yml`), the cleanest path is to remove the old image first, then
rerun `ensure_image`:

```
docker image rm $KERMT_IMAGE
$HELPER ensure_image
```

Confirm with the user before running `docker image rm`.
molmim-nim8.23 KB

View saved version →

---
name: molmim-nim
description: >
  Use this skill for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization. Invoke for MolMIM, molecular embeddings, hidden states, latent decoding, sampling around a seed SMILES, CMA-ES guided molecule generation, QED or plogP optimization, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28; rdkit"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# MolMIM NIM

Generate, sample, embed, and decode small molecules with MolMIM. Use this
`SKILL.md` for first-pass hosted/local usage; load supplemental files only when
needed:

- `references/api.md`: endpoints, schema, Docker flags, response fields.
- `references/science.md`: use cases, strengths, limits, and handoffs.
- `references/parameters.md`: generation, sampling, and optimization effects.
- `references/validation.md`: SMILES/property/artifact checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted generation: `https://health.api.nvidia.com/v1/biology/nvidia/molmim/generate`
- Local generation: `http://localhost:8000/generate`
- Local embedding: `http://localhost:8000/embedding`
- Local hidden state: `http://localhost:8000/hidden`
- Local decode: `http://localhost:8000/decode`
- Local sampling: `http://localhost:8000/sampling`
- Local readiness: `http://localhost:8000/v1/health/ready`

Mode difference: the hosted API reference exposes `/generate`; the local
container exposes the broader latent-space workflow (`/embedding`, `/hidden`,
`/decode`, `/sampling`, `/generate`). Do not invent hosted latent endpoints.

Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Local inference uses
no auth header after readiness.

## Local Docker

Use shell env first; source repo-root `.env` only if present. Do not print keys.
MolMIM docs use `NGC_CLI_API_KEY` for the local container; this repo accepts
`NGC_API_KEY` or `NVIDIA_API_KEY` and maps to `NGC_CLI_API_KEY` for startup.
Mount `LOCAL_NIM_CACHE` at `/home/nvs/.cache/nim`.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
if [ -z "${NGC_CLI_API_KEY:-}" ] && [ -n "${NGC_API_KEY:-}" ]; then
  export NGC_CLI_API_KEY="$NGC_API_KEY"
fi
: "${NGC_CLI_API_KEY:?Set NGC_API_KEY, NVIDIA_API_KEY, or NGC_CLI_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_CLI_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm -it --name molmim \
  --runtime=nvidia \
  -e CUDA_VISIBLE_DEVICES="${NIM_TEST_GPU}" \
  -e NGC_CLI_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/home/nvs/.cache/nim" \
  -p 8000:8000 \
  nvcr.io/nim/nvidia/molmim:1.0.0
```

Readiness check:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

Local embedding smoke test after readiness. Local inference uses no
`Authorization` header:

```python
import requests

seed = "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
response = requests.post(
    "http://localhost:8000/embedding",
    headers={"Content-Type": "application/json"},
    json={"sequences": [seed]},
    timeout=60,
)
response.raise_for_status()
embedding_data = response.json()
embeddings = embedding_data["embeddings"]
print(f"received {len(embeddings)} embedding vector(s)")
```

## Hosted Generation Pattern

Use hosted `/generate` for seed-SMILES generation or optimization. Use
`algorithm: "CMA-ES"` for guided property optimization and `algorithm: "none"`
for unguided sampling around the seed.

```python
import os
import requests

hosted = True
url = (
    "https://health.api.nvidia.com/v1/biology/nvidia/molmim/generate"
    if hosted else "http://localhost:8000/generate"
)
headers = {"Content-Type": "application/json"}
if hosted:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "smi": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C",
    "algorithm": "CMA-ES",
    "num_molecules": 10,
    "property_name": "QED",
    "minimize": False,
    "min_similarity": 0.4,
    "particles": 8,
    "iterations": 3,
}

response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
```

Generation gotchas:

- Field name is `smi`, not `smiles`.
- `algorithm` is `"CMA-ES"` or `"none"`.
- `property_name` is `"QED"` or `"plogP"`.
- `num_molecules` is 1-100. `iterations` is 1-1000. `particles` is 2-1000.
- `min_similarity` is 0-1 in the hosted API reference; local docs emphasize
  common values up to 0.7 for constrained optimization.
- `scaled_radius` is 0-2 and is mainly used with `algorithm: "none"` or local
  `/sampling`.

## Local Latent Workflow

Use local-only endpoints for embedding, hidden-state manipulation, and decode.
This is also the surface used by the guided optimization example package.
For local latent workflows, state explicitly that the hosted API reference
exposes `/generate`; `/embedding`, `/hidden`, `/decode`, and `/sampling` are
local-only in the current docs.

```python
seed = "CC(Cc1ccc(cc1)C(C(=O)O)C)C"
base = "http://localhost:8000"
headers = {"Content-Type": "application/json"}

embedding = requests.post(
    f"{base}/embedding",
    headers=headers,
    json={"sequences": [seed]},
    timeout=60,
)
embedding.raise_for_status()
embedding_data = embedding.json()
embeddings = embedding_data["embeddings"]
print(f"received {len(embeddings)} embedding vector(s)")

hidden = requests.post(
    f"{base}/hidden",
    headers=headers,
    json={"sequences": [seed]},
    timeout=60,
)
hidden.raise_for_status()
hidden_data = hidden.json()
hiddens = hidden_data["hiddens"]
mask = hidden_data["mask"]

decoded = requests.post(
    f"{base}/decode",
    headers=headers,
    json={"hiddens": hiddens, "mask": mask},
    timeout=60,
)
decoded.raise_for_status()

sampled = requests.post(
    f"{base}/sampling",
    headers=headers,
    json={"sequences": [seed], "num_molecules": 10, "scaled_radius": 0.7},
    timeout=60,
)
sampled.raise_for_status()
```

## Save And Validate Output

Save generated SMILES and validate before using them downstream.

```python
from pathlib import Path
import json

def molmim_smiles(result):
    values = []
    if isinstance(result.get("generated"), list):
        for item in result["generated"]:
            if isinstance(item, str):
                values.append(item)
            elif isinstance(item, list):
                values.extend(x for x in item if isinstance(x, str))
    molecules = result.get("molecules")
    if isinstance(molecules, str):
        molecules = json.loads(molecules)
    if isinstance(molecules, list):
        for item in molecules:
            if isinstance(item, dict) and isinstance(item.get("sample"), str):
                values.append(item["sample"])
    return values

generated = molmim_smiles(result)
if not generated:
    raise RuntimeError(f"MolMIM returned no generated molecules: {result}")

Path("molmim_response.json").write_text(json.dumps(result, indent=2))
Path("molmim_generated.smi").write_text("\n".join(generated) + "\n")
for i, smiles in enumerate(generated, start=1):
    print(i, smiles)
```

Use RDKit when available to check parseability, uniqueness, simple property
ranges, and whether seed similarity constraints are plausible. Generated
molecules are candidates, not validated hits; use downstream property, docking,
affinity, toxicity, and synthetic-feasibility checks before prioritization.

## Troubleshooting

- Hosted `404` on `/embedding`, `/hidden`, `/decode`, or `/sampling`: those
  endpoints are local-only in the docs.
- `401`: missing or unauthorized NGC key for hosted requests.
- Hosted response parsing: live hosted `/generate` may return `molecules` as
  a JSON string of `{sample, score}` objects, while local endpoints may return
  `generated`; parse both.
- `422`: invalid SMILES, unsupported `algorithm`, invalid `property_name`, or
  parameter outside documented ranges.
- Local startup auth: set `NGC_CLI_API_KEY`, or set `NGC_API_KEY`/`NVIDIA_API_KEY`
  and map it as shown above.
- Local startup cache misses: mount `LOCAL_NIM_CACHE` to `/home/nvs/.cache/nim`,
  not `/opt/nim/.cache`.

Referenced files: 5

msa-search-nim16.4 KB

View saved version →

---
name: msa-search-nim
description: >
  Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local Docker deployment. For local deployment, download the databases in parallel with aria2c and launch via NIM_MODEL_NAME (the recommended default fast path, ~14 min vs >80 min for the built-in downloader); a plain docker run uses the slow built-in downloader.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# MSA-Search NIM

Generate protein MSAs with GPU-accelerated MMSeqs2. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: MSA purpose, pairing/templates, limits, handoffs.
- `references/parameters.md`: database, pairing, depth, and template tuning.
- `references/validation.md`: alignment, template, and artifact checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode And Endpoint

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted standard MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict`
- Hosted paired MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict`
- Local standard MSA: `http://localhost:8000/biology/colabfold/msa-search/predict`
- Local paired MSA: `http://localhost:8000/biology/colabfold/msa-search/paired/predict`
- Local templates: `http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict`

Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
The hosted template path returned HTTP 404 in validation, so use local Docker
for template search unless the hosted docs/service changes.

## Local Docker

**Default local deployment = parallel download + `NIM_MODEL_NAME`.** The first recipe below
is the one to use for real workflows. It downloads the database(s) with a range-parallel
downloader (aria2c) and starts the NIM against those files — ~14 min for UniRef30 vs >80 min
for the NIM's built-in downloader (measured, H100). Do **not** reach for the plain `docker run`
(the "Fallback" subsection) unless you only want a `databases:pdb70` smoke test or you
deliberately want the NIM to manage its own blob cache.

Local setup requires a GPU. Size the NVMe volume to the profile you pick (UniRef30 ~490 GB;
full set ~1.4 TB). For setup answers, include env preflight, `docker login`, the parallel
download, `NIM_MODEL_NAME` launch, readiness, and then no-auth local inference. Do not invent a
cache default or drop the `NVIDIA_API_KEY` fallback.

```bash
# --- env preflight (do not drop the NVIDIA_API_KEY fallback) ---
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${DB_DIR:=/data/fast-db}"          # where the parallel download lands

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

# --- 1) pick the DB version(s) you need (paired/complex work = uniref30 only) ---
DB_VERSION=uniref30_2302-m18v1
command -v aria2c >/dev/null || sudo apt-get install -y aria2
mkdir -p "$DB_DIR"

# --- 2) parallel download from NGC (see "Parallel Download" section for the all-DB loop) ---
curl -fsS -H "Authorization: Bearer $NGC_API_KEY" \
  "https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/${DB_VERSION}/files" \
  -o /tmp/files.json
DB_DIR="$DB_DIR" python3 - <<'PY'
import json, os
d = json.load(open("/tmp/files.json")); dbdir = os.environ["DB_DIR"]; lines = []
for url, path in zip(d["urls"], d["filepath"]):
    lines += [url.strip(), f"  dir={dbdir}", f"  out={path}"]
open("/tmp/aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i /tmp/aria.in --max-concurrent-downloads=4 --max-connection-per-server=16 \
  --split=16 --min-split-size=1M --continue=true --file-allocation=none

# --- 3) launch the NIM against the downloaded files (skips the slow built-in download) ---
docker run -d --name msa-search --runtime=nvidia --gpus all \
  -e NGC_API_KEY \
  -e NIM_MODEL_NAME=/databases \
  -v "${DB_DIR}:/databases" \
  -p 8000:8000 \
  nvcr.io/nim/colabfold/msa-search:2
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

If the DB is already present in `$DB_DIR`, skip steps 1-2 — the launch alone is a ~20 s warm
start. See "Parallel Download For Any Database Set" for the multi-database (`databases:all`)
loop and the full rationale.

### Fallback: Let The NIM Download Its Own Databases (slower)

Use this only for a quick `databases:pdb70` smoke test, or when you specifically want the NIM
to manage its own blob cache. It uses the built-in downloader, which is slow on large profiles
(UniRef30 stalled past 80 min in testing). Pin the smallest profile with `NIM_MODEL_PROFILE`
(see "Faster Startup") so it does not fetch the full 1.4 TB.

```bash
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
mkdir -p "${LOCAL_NIM_CACHE}"; chmod 777 "${LOCAL_NIM_CACHE}"
docker run --rm --name msa-search \
  --runtime=nvidia --gpus all \
  -e NGC_API_KEY \
  -e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/colabfold/msa-search:2
```


## Faster Startup: Task-Specific Database Profiles

The full database download is ~1.4 TB and can take well over an hour on first launch. If you
only need some databases, select a **task-specific profile** so the NIM downloads just those.
This is the single biggest lever on local startup time.

List the profiles your image actually ships (hashes change between releases — never hardcode
them):

```bash
docker run --rm --entrypoint list-model-profiles nvcr.io/nim/colabfold/msa-search:2
```

Then pass the chosen hash with `NIM_MODEL_PROFILE`:

```bash
docker run --rm --name msa-search \
  --runtime=nvidia --gpus all \
  -e NGC_API_KEY \
  -e NIM_MODEL_PROFILE=<hash-from-list-model-profiles> \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/colabfold/msa-search:2
```

Profiles available in this image (confirm hashes with `list-model-profiles`):

| Profile tags | Databases | Best for | Storage |
|---|---|---|---|
| `databases:pdb70` | PDB70 | Quick testing / smoke check | ~100 MB |
| `databases:uniref30` | UniRef30 | **Paired MSA search for complexes** — UniRef30 is the only DB used for species-based pairing | ~500 GB |
| `databases:uniref30,pdb70,pdb` | UniRef30 + PDB70 + PDB structures | Structural template search | ~700 GB |
| `databases:all` (default) | UniRef30 + ColabFold envdb + PDB70 + PDB100 + PDB structures | Full sensitivity, all databases | ~1.2 TB |

Verify the loaded profile after readiness:

```bash
curl -s localhost:8000/v1/metadata | jq
```

Notes:

- The request-level `databases` parameter only selects among databases **already
  downloaded**; it does NOT change what is fetched at startup. Startup footprint is set by
  `NIM_MODEL_PROFILE` alone.
- **Paired search needs UniRef30 only.** `colabfold_envdb_202108` has no taxonomy and cannot
  be used for pairing, so `databases:uniref30` is the correct, smallest profile for
  complex/paired workflows — it skips the envdb, the largest part of the full set.
- For maximum monomer sensitivity (UniRef30 + envdb merged) you still need `databases:all`;
  there is no envdb-inclusive profile smaller than the full set.

### Custom Or Individual Databases

To use a single manually downloaded database (or your own MMSeqs2 DB), download it from NGC
and point the NIM at the mount with `NIM_MODEL_NAME` instead of a profile:

```bash
ngc registry model download-version nim/colabfold/msa-search:uniref30_2302-m18v1
# then mount the directory and set -e NIM_MODEL_NAME=/databases
```

`NIM_MODEL_NAME` **replaces** the profile databases entirely — the NIM uses only what is
under that directory (discovered by scanning for `**/*.idx`). Mount multiple databases under
one parent to combine them. NGC-downloaded databases are pre-indexed for GPU Server; custom
databases must be indexed with `mmseqs createindex` first. Individually downloadable NGC model
versions: `uniref30_2302-m18v1`, `colabfold_envdb_202108-m18v1`, `pdb70_220313-m18v1`,
`pdb100_230517-m18v1`, `pdb_20251028_zip-m18v1`.

## Recommended: Parallel Download For Any Database Set (Fast Deployment)

**This is the recommended way to download the databases at all — for any profile,
including the full `databases:all` set.** Task-specific profiles cut *what* you
download; this parallel downloader cuts *how long* that download takes. Use it whether
you need one database or all of them. The gain is largest for the `databases:uniref30` profile, which is ~490 GB
dominated by two very large files (a ~241 GB GPU index and a ~134 GB sequence DB).

The NIM's built-in downloader parallelizes **across files** (`max_parallel_files=10`) but pulls
each file over roughly one connection. The NGC CDN throttles a single connection to ~20–25
MB/s, so while the downloader is fetching one of the two giant files, most of its parallel
slots sit idle and throughput collapses to that single-flow rate. Measured on an H100 node,
the built-in path did not reach `/health/ready` in over 80 minutes.

A range-parallel downloader splits **each file** into many byte-range segments (the NGC CDN
advertises `accept-ranges: bytes`), so a single 241 GB file is pulled over 16 connections at
once — ~15× the single-flow rate. Same node, `aria2c` fetched the full ~490 GB in **~13.5
minutes**.

Workflow (download once with aria2, then start the NIM against the files via `NIM_MODEL_NAME`):

```bash
# 1) Get presigned file URLs for the individual database model version from NGC.
#    (Requires NGC_API_KEY. The response arrays `urls` and `filepath` are positionally paired.)
curl -s -H "Authorization: Bearer $NGC_API_KEY" \
  'https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/uniref30_2302-m18v1/files' \
  -o files.json

# 2) Build an aria2 input file (URL + target filename per entry) and download in parallel.
python3 - <<'PY'
import json
d = json.load(open("files.json"))
lines = []
for url, path in zip(d["urls"], d["filepath"]):
    lines += [url.strip(), "  dir=/data/fast-db", f"  out={path}"]
open("aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i aria.in \
  --max-concurrent-downloads=4 --max-connection-per-server=16 --split=16 \
  --min-split-size=1M --continue=true --file-allocation=none

# 3) Start the NIM against the downloaded directory. NIM_MODEL_NAME makes the NIM discover
#    databases by scanning for **/*.idx, bypassing the profile/blob cache entirely.
docker run -d --name msa-search --runtime=nvidia --gpus all \
  -e NGC_API_KEY \
  -e NIM_MODEL_NAME=/databases \
  -v /data/fast-db:/databases \
  -p 8000:8000 \
  nvcr.io/nim/colabfold/msa-search:2
```

For **all databases** (equivalent to `databases:all`), repeat step 1 for each individual DB
version and download them into sibling directories under one parent, then point
`NIM_MODEL_NAME` at that parent — the NIM discovers every DB by scanning `**/*.idx`:

```bash
# fetch each DB's file list into /data/all-db/<db>/ ... then one aria2c per list, e.g.:
for V in uniref30_2302-m18v1 colabfold_envdb_202108-m18v1 pdb70_220313-m18v1 \
         pdb100_230517-m18v1 pdb_20251028_zip-m18v1; do
  curl -s -H "Authorization: Bearer $NGC_API_KEY" \
    "https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/$V/files" \
    -o "files_$V.json"
  # build an aria2 input from files_$V.json (dir=/data/all-db) and run aria2c on it
done
# then launch once against the parent:
#   docker run -d ... -e NIM_MODEL_NAME=/databases -v /data/all-db:/databases ...
```

The per-connection CDN throttle is the same for every database, so parallel download helps
the full set proportionally — the more you download, the more absolute time it saves.

Notes:

- The presigned URLs expire (typically within a day) — build the aria2 input and start the
  download promptly after fetching `files.json`.
- Keep the downloaded directory's internal layout intact (e.g. `uniref30_2302/…`); the
  `filepath` values already encode it. The NIM needs the `.idx` file plus its companion files
  and the small `.UNIREF30_READY` / `*.tar.gz.unpacked` markers.
- The bottleneck is the CDN's per-connection cap, not local disk or CPU — a fast NVMe volume
  writes far faster than the network delivers. Raising `--split` / `--max-connection-per-server`
  helps only up to the node's aggregate egress ceiling.
- Best of all: download the profile once, then **persist the cache volume** (or this
  `fast-db` directory) and mount it on future nodes for a ~20 s warm start with no re-download.

## Standard MSA Request

Use exact case-sensitive database names and response keys.

```python
import os
import requests

HOSTED = True
url = (
    "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict"
    if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "sequence": "SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR",
    "databases": ["Uniref30_2302", "colabfold_envdb_202108"],
    "e_value": 0.0001,
    "output_alignment_formats": ["a3m"],
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

## Paired MSA Request

Use paired search for protein complexes; payload field is `sequences` plural,
and output is `alignments_by_chain`.

```python
url = (
    "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict"
    if HOSTED else "http://localhost:8000/biology/colabfold/msa-search/paired/predict"
)
payload = {
    "sequences": [chain_a_sequence, chain_b_sequence],
    "e_value": 0.0001,
    "output_alignment_formats": ["a3m"],
}
```

## Local Template Search

Use local Docker for structural templates. Set `max_msa_sequences=500` unless
`NIM_GLOBAL_MAX_MSA_DEPTH` was changed.

```python
url = "http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict"
headers = {"Content-Type": "application/json"}
payload = {
    "sequence": "VLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHGKKVADALTNAVA",
    "structural_template_databases": ["pdb70_220313"],
    "max_structures": 20,
    "max_msa_sequences": 500,
}
```

## Save Outputs

```python
# Standard MSA: result["alignments"][database][format]["alignment"]
for db_name, formats in result.get("alignments", {}).items():
    for fmt_name, data in formats.items():
        with open(f"msa_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
            handle.write(data["alignment"])

# Paired MSA: one alignment set per chain
for chain_id, chain_data in result.get("alignments_by_chain", {}).items():
    for db_name, formats in chain_data.items():
        for fmt_name, data in formats.items():
            with open(f"msa_chain_{chain_id}_{db_name}.{fmt_name}", "w", encoding="utf-8") as handle:
                handle.write(data["alignment"])

# Template search: save mmCIF structures and M8 hit tables
for name, cif in result.get("structures", {}).items():
    open(f"template_{name}.cif", "w", encoding="utf-8").write(cif)
for name, hit_table in result.get("search_hits", {}).items():
    open(f"template_hits_{name}.m8", "w", encoding="utf-8").write(hit_table)
```

A3M output can feed OpenFold3, AlphaFold2, or RoseTTAFold. For alignment depth,
template, and sequence sanity checks, read `references/validation.md`.

## Limits And Troubleshooting

- Sequence length: 1-4096 amino acids; `X` works since v2.3.0.
- `max_msa_sequences`: 1-500; local GPU server default must match
  `NIM_GLOBAL_MAX_MSA_DEPTH`.
- Paired MSA requires at least two sequences.
- Local URL 404 usually means an accidental `/v1/` prefix.
- First local run can take hours while databases populate `LOCAL_NIM_CACHE`.

Referenced files: 5

msa-structure-prediction-pipeline6.2 KB

View saved version →

---
name: msa-structure-prediction-pipeline
description: >
  Run a complete protein structure prediction pipeline using NVIDIA BioNeMo NIMs:
  search for MSA alignments with MSA-Search (ColabFold), then predict the structure
  with OpenFold3 using the retrieved alignments. Use this skill whenever the user wants
  to predict a protein structure with maximum accuracy using MSA context, run the
  full AlphaFold3-style pipeline, generate MSA-informed structure predictions, or
  improve structure prediction accuracy by providing evolutionary information.
  Triggers on: MSA structure prediction pipeline, structure prediction pipeline, MSA-informed prediction, OpenFold3,
  ColabFold MSA, AlphaFold3 pipeline, protein structure, homology search, a3m alignment,
  UniRef30, NIM microservice. This pipeline chains MSA-Search and OpenFold3.
license: Apache-2.0 AND CC-BY-4.0
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# MSA Structure Prediction Pipeline

Predict protein structures with high accuracy by chaining two BioNeMo NIMs:

```
Step 1: MSA-Search  →  Step 2: OpenFold3
(Search homologs)       (Predict structure with MSA)
```

---

## Overview

Why chain these NIMs?
- **MSA-Search** finds evolutionary homologs in UniRef30 and ColabFold databases using GPU-accelerated MMSeqs2. The resulting alignment provides crucial evolutionary information.
- **OpenFold3** uses the MSA to improve structure prediction accuracy — especially for sequences where no close homolog exists in PDB.
- Running MSA-Search first means OpenFold3 gets the full evolutionary context rather than a single-sequence prediction.

---

## Before you start

Confirm with the user:
1. **Query sequence**: amino acid sequence to predict
2. **MSA depth**: how many sequences to retrieve (default 500; more = slower but more context)
3. **API mode**: hosted or local Docker?

Note: local MSA-Search requires 1.4 TB of database storage — strongly recommend hosted unless the user has that infrastructure.

For local Docker, do not assume MSA-Search and OpenFold3 are both on
`localhost:8000` concurrently. Run one container at a time and hand off the A3M
file, or start each NIM on a distinct host port and set the URLs explicitly.

---

## Step 1: Search for MSA with MSA-Search

```python
import requests, json, os
from pathlib import Path

NGC_API_KEY = os.environ["NGC_API_KEY"]
HOSTED = True

query_sequence = "<YOUR_PROTEIN_SEQUENCE>"

if HOSTED:
    msa_url = "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict"
    headers = {"Content-Type": "application/json",
               "Authorization": f"Bearer {NGC_API_KEY}"}
else:
    msa_url = "http://localhost:8000/biology/colabfold/msa-search/predict"
    headers = {"Content-Type": "application/json"}

payload = {
    "sequence": query_sequence,
    "databases": ["Uniref30_2302", "colabfold_envdb_202108"],
    "e_value": 0.0001,
    "output_alignment_formats": ["a3m"],
}

r = requests.post(msa_url, headers=headers, json=payload)
r.raise_for_status()
msa_result = r.json()

# Extract the A3M alignment
a3m_alignment = msa_result["alignments"]["Uniref30_2302"]["a3m"]["alignment"]

# Save for reference
with open("query_msa.a3m", "w") as f:
    f.write(a3m_alignment)

# Count sequences in alignment
n_seqs = a3m_alignment.count(">")
print(f"Step 1 complete: found {n_seqs} homologous sequences")
print(f"MSA saved to query_msa.a3m")
```

---

## Step 2: Predict structure with OpenFold3

Pass the MSA directly into OpenFold3's `msa` field:

```python
if HOSTED:
    of3_url = "https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict"
else:
    of3_url = "http://localhost:8000/biology/openfold/openfold3/predict"

# Build the OpenFold3 MSA structure from the retrieved alignment
msa_data = {
    "uniref30": {
        "a3m": {
            "alignment": a3m_alignment,
            "format": "a3m"
        }
    }
}

# Optionally also include colabfold_envdb alignment if requested
# env_alignment = msa_result["alignments"]["colabfold_envdb"]["a3m"]["alignment"]
# msa_data["colabfold_env"] = {"a3m": {"alignment": env_alignment, "format": "a3m"}}

payload = {
    "inputs": [{
        "input_id": "prediction_with_msa",
        "output_format": "pdb",
        "molecules": [
            {
                "type": "protein",
                "sequence": query_sequence,
                "diffusion_samples": 1,
                "msa": msa_data
            }
        ]
    }]
}

r = requests.post(of3_url, headers=headers, json=payload, timeout=300)
r.raise_for_status()
result = r.json()

output = result["outputs"][0]
for i, sample in enumerate(output["structures_with_scores"]):
    fmt = sample["format"]
    filename = f"predicted_structure_{i+1}.{fmt}"
    with open(filename, "w") as f:
        f.write(sample["structure"])
    print(f"\nStep 2 complete: {filename} saved")
    print(f"  Confidence:  {sample['confidence_score']:.4f}")
    print(f"  pLDDT:       {sample['complex_plddt_score']:.4f}")
    print(f"  pTM:         {sample['ptm_score']:.4f}")
```

---

## Comparing single-sequence vs MSA-informed prediction

If the user wants to see the impact of MSA, run OpenFold3 twice — once with the full MSA and once with just the query sequence as a minimal alignment:

```python
# Minimal MSA (single sequence — same as no MSA context):
minimal_msa = {
    "main": {
        "a3m": {
            "alignment": f">query\n{query_sequence}",
            "format": "a3m"
        }
    }
}
```

A larger, higher-quality MSA typically yields higher pLDDT and lower pDE, especially for proteins with many known homologs.

---

## For protein complexes

Use the `/paired/predict` endpoint of MSA-Search to get paired alignments for multi-chain complexes, then pass each chain's alignment into the corresponding molecule's `msa` field and `paired_msa` fields:

```python
# Paired MSA search endpoint for complexes:
msa_paired_url = "https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict"
paired_payload = {
    "sequences": [chain_A_sequence, chain_B_sequence],
    "e_value": 0.0001,
}
```

---

## Quick reference — skill dependencies

| Step | Skill | Key endpoint |
|---|---|---|
| MSA search | `msa-search-nim` | `/biology/colabfold/msa-search/predict` |
| Structure prediction | `openfold3-nim` | `/biology/openfold/openfold3/predict` |
nvmolkit-usage17.3 KB

View saved version →

---
name: nvmolkit-usage
description: >-
  Write code that calls the installed nvMolKit Python API for GPU-accelerated,
  batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine
  similarity, ETKDG conformer embedding, MMFF/UFF optimization, TFD, conformer
  RMSD, Butina clustering, and substructure search. Use when the user is
  importing `nvmolkit.*`, debugging an `nvmolkit` call, choosing between
  nvMolKit and RDKit for a batched cheminformatics workflow, or wiring nvMolKit
  results into a torch/numpy pipeline. Out of scope: building nvMolKit from
  source.
license: Apache-2.0
metadata:
  owner: Kevin Boyd (@scal444)
  risk_tier: skill
---

# nvMolKit usage

## What nvMolKit is

GPU-accelerated, batched implementations of common RDKit operations. APIs mirror RDKit where possible but are batch-oriented: they take lists of `rdkit.Chem.Mol` (or lists of fingerprints) and process them in parallel on one or more GPUs. nvMolKit links against RDKit at build time; inputs and outputs are real RDKit `Mol` objects.

## Where nvMolKit does well

Reach for nvMolKit when:

- The workload is **a large batch of molecules** processed together (typically thousands or more).
- The metric is **throughput / total wall time across the batch**, not per-molecule latency.
- The same operation is **repeated identically** across the batch (fingerprinting a library, embedding/minimizing many conformers, bulk pairwise similarity), so the GPU stays saturated.

Plain RDKit is usually the better choice for single-molecule one-offs or workflows that can't be expressed as a batch. nvMolKit is not meant to replace RDKit for those cases.

## Runtime requirements

- An NVIDIA GPU with compute capability 7.0 (V100) or higher
- A CUDA driver compatible with CUDA 12.6+.
- A working `torch` install with CUDA support (nvMolKit returns GPU tensors via `torch`'s CUDA array interface).

If CUDA is unavailable, nvMolKit calls raise. There is no CPU fallback - if the user needs one, use RDKit directly for that path.

When helping with installation, make the user choose a PyTorch CUDA backend that the host driver supports before installing nvMolKit. nvMolKit's PyPI wheels are built with CUDA Toolkit 12.9 and depend on CUDA 12 runtime packages, but pip/uv can still select a CUDA 13 PyTorch wheel unless the install command says otherwise.

- Conda: prefer conda-forge `pytorch-gpu`; pin `cuda-version=12.6` or another CUDA version supported by the driver.
- pip: send the user to the PyTorch install selector (`https://pytorch.org/get-started/locally/`) or previous-versions page (`https://pytorch.org/get-started/previous-versions/`) to install `torch` for a CUDA 12.x backend before installing nvMolKit.
- uv: install nvMolKit with an explicit backend, e.g. `uv pip install --torch-backend=cu128 nvmolkit`.

## Verify the install before writing real code

Run this once to confirm nvMolKit is importable and a GPU op works end to end:

```python
import nvmolkit
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator

print("nvmolkit:", nvmolkit.__version__)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())

mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
result = fpgen.GetFingerprints(mols)
torch.cuda.synchronize()
fps = result.torch()
print("fps shape:", tuple(fps.shape), "dtype:", fps.dtype)
# Expected: shape (3, 32), dtype torch.int32  (1024 bits packed into 32 int32s per row)
```

If this fails, point the user at the install guide on the docs site rather than guessing - see "Going deeper" below.

## Entry points

| Task | Module | Primary entry point |
|---|---|---|
| Morgan fingerprints | `nvmolkit.fingerprints` | `MorganFingerprintGenerator(radius, fpSize).GetFingerprints(mols)` |
| Bulk Tanimoto / cosine similarity | `nvmolkit.similarity` | `crossTanimotoSimilarity(...)`, `crossCosineSimilarity(...)`, plus `*MemoryConstrained` variants for results too large to fit in GPU memory |
| ETKDG conformer embedding | `nvmolkit.embedMolecules` | `EmbedMolecules(molecules, params, confsPerMolecule, ...)` |
| MMFF94 optimization (one-shot) | `nvmolkit.mmffOptimization` | `MMFFOptimizeMoleculesConfs(molecules, ...)` |
| UFF optimization (one-shot) | `nvmolkit.uffOptimization` | `UFFOptimizeMoleculesConfs(molecules, ...)` |
| Forcefield with custom options + constraints | `nvmolkit.batchedForcefield` | `MMFFBatchedForcefield(mols, properties=..., nonBondedThreshold=..., ignoreInterfragInteractions=..., hardwareOptions=...)`, `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`. Per-molecule view `ff[i]` exposes `add_distance_constraint`, `add_position_constraint`, `add_angle_constraint`, `add_torsion_constraint`. Methods: `.compute_energy()`, `.compute_gradients()`, `.minimize(maxIters, forceTol)` |
| Pairwise conformer RMSD | `nvmolkit.conformerRmsd` | `GetConformerRMSMatrix(mol)`, `GetConformerRMSMatrixBatch(mols)` |
| Torsion Fingerprint Deviation (TFD) | `nvmolkit.tfd` | `GetTFDMatrix(mol)`, `GetTFDMatrices(mols)` |
| Butina clustering | `nvmolkit.clustering` | `butina(distance_matrix, cutoff)` (precomputed matrix), `fused_butina(fingerprints, cutoff)` (memory-efficient, on-the-fly) |
| Substructure search | `nvmolkit.substructure` | `hasSubstructMatch`, `countSubstructMatches`, `getSubstructMatches` |
| Hardware tuning (batch size, GPU IDs) | `nvmolkit.types` | `HardwareOptions(...)` passed to ETKDG / MMFF / UFF |
| Optional autotuning of `HardwareOptions` | `nvmolkit.autotune` | `tune_embed_molecules`, `tune_mmff_optimize`, `tune_uff_optimize`, `tune_batched_forcefield`. Requires the `optuna` package |

## Result types and execution model

Two return shapes carry GPU-resident output, depending on what the operation produces.

### `AsyncGpuResult`

Used by operations that return a single flat tensor (fingerprints, similarity matrices, RMSD/TFD vectors, Butina inputs). Key behaviors:

- Asynchronous. The kernel may not have completed when the call returns.
- `result.torch()` returns a zero-copy `torch.Tensor` on the GPU. Caller is responsible for synchronizing before reading values on the host.
- `result.numpy()` synchronizes and returns a CPU numpy array.
- Exposes `__cuda_array_interface__`, so it can be passed directly into other nvMolKit functions (e.g. fingerprints → similarity) with no host round-trip.

#### CUDA stream control

A subset of the `AsyncGpuResult`-returning APIs accept an optional `stream: torch.cuda.Stream | None = None` argument so callers can submit nvMolKit work to a non-default stream and overlap it with their own kernels. When omitted, the call uses the current torch stream.

APIs that take a `stream` argument:

- `MorganFingerprintGenerator.GetFingerprints`
- `crossTanimotoSimilarity`, `crossCosineSimilarity`, and their `*MemoryConstrained` variants
- `butina`, `fused_butina`
- `GetConformerRMSMatrix`, `GetConformerRMSMatrixBatch`

Other APIs (ETKDG, MMFF/UFF optimization, TFD, substructure search) are synchronous to the caller — no stream plumbing needed.

Typical pattern:

```python
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity

stream = torch.cuda.Stream()
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]

with torch.cuda.stream(stream):
    fps = fpgen.GetFingerprints(mols, stream=stream)
    sim = crossTanimotoSimilarity(fps, stream=stream)
stream.synchronize()
print(sim.torch())
```

### `Device3DResult`

Used by ETKDG embedding and MMFF/UFF optimization (one-shot and `BatchedForcefield`) when called with `output=CoordinateOutput.DEVICE`. The GPU-resident equivalent of writing conformers back to `Mol` objects. Fields:

- `values`: `AsyncGpuResult` of shape `(total_atoms, 3)` float64. Concatenated conformer coordinates in CSR-style layout.
- `atom_starts`, `mol_indices`, `conf_indices`: `AsyncGpuResult` int32 buffers describing the layout (`values[atom_starts[i]:atom_starts[i+1]]` is conformer `i`'s atoms).
- `energies`, `converged`: `AsyncGpuResult` buffers populated only for MMFF/UFF minimization (not for plain ETKDG).
- `gpu_id`: device the buffers live on. The `targetGpu` argument on each API picks this; `targetGpu=-1` uses the default consolidation device.
- `.per_molecule()` returns nested `list[list[torch.Tensor]]` of per-conformer views; `.dense(pad_value=nan)` materializes a padded `(n_mols, max_confs, max_atoms, 3)` tensor.

The default mode (`CoordinateOutput.RDKIT_CONFORMERS`) still writes optimized coordinates back into each `Mol` and returns Python lists of energies/convergence flags. Reach for `CoordinateOutput.DEVICE` when chaining downstream GPU work (e.g. ETKDG → MMFF → similarity scoring) without host round-trips.

## Configuration

Two configuration objects expose the GPU/CPU knobs.

### `HardwareOptions` (ETKDG, MMFF, UFF)

`from nvmolkit.types import HardwareOptions`. Passed via `hardwareOptions=` to `EmbedMolecules`, `MMFFOptimizeMoleculesConfs`, `UFFOptimizeMoleculesConfs`, and the `BatchedForcefield` constructors. Every field has an "auto" sentinel; the defaults are usually fine.

| Field | Type | Default | Meaning |
|---|---|---|---|
| `preprocessingThreads` | int | `-1` (all visible CPUs) | CPU threads for preprocessing |
| `batchSize` | int | `-1` (auto-tuned) | Number of conformers per GPU batch |
| `batchesPerGpu` | int | `-1` (auto) | Concurrent batches per GPU; must be `>0` or `-1` |
| `gpuIds` | `list[int]` | `[]` (all visible GPUs) | Specific device ordinals to target |

Passing a `gpuIds` entry for a device that isn't visible raises `RuntimeError: invalid device ordinal`. For finding good values automatically across a representative sample, see `nvmolkit.autotune` (requires the `optuna` extra); each `tune_*` function returns a `TuneResult` whose `best_config` is a fully-populated `HardwareOptions` ready to pass back into the real call.

`HardwareOptions` round-trips through `to_dict()` / `from_dict()` for persisting tuned configs to disk.

### `SubstructSearchConfig` (substructure search)

`from nvmolkit.substructure import SubstructSearchConfig`. Passed via `config=` to `hasSubstructMatch`, `countSubstructMatches`, and `getSubstructMatches`.

| Field | Type | Default | Meaning |
|---|---|---|---|
| `batchSize` | int | `1024` | (target, query) pairs per GPU batch |
| `workerThreads` | int | `-1` (auto) | GPU runner threads per GPU |
| `preprocessingThreads` | int | `-1` (auto) | CPU threads for preprocessing |
| `maxMatches` | int | `0` (unlimited) | Max matches returned per (target, query) pair |
| `uniquify` | bool | `False` | Drop duplicate matches that differ only in atom enumeration order |
| `gpuIds` | `list[int] \| None` | `None` (current device only) | Specific device ordinals to target |

Substructure search currently does not support chirality-aware matching, enhanced stereochemistry, or other advanced RDKit `SubstructMatchParameters` options.

## Recipes

### Morgan fingerprints + bulk Tanimoto similarity

```python
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity

smiles = ["CCO", "CCN", "c1ccccc1", "CC(=O)O", "CCOCC"]
mols = [Chem.MolFromSmiles(smi) for smi in smiles]

fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
fps = fpgen.GetFingerprints(mols)

sim = crossTanimotoSimilarity(fps)
torch.cuda.synchronize()
print(sim.torch())
```

Inputs are `list[Mol]`. Output of `GetFingerprints` is an `AsyncGpuResult` wrapping an `(n_mols, fpSize / 32)` int32 tensor of packed bits. Pass it straight into `crossTanimotoSimilarity` for an `(n, n)` similarity matrix; pass two fingerprint sets for an `(n, m)` cross-matrix. For sets too large to materialize on the GPU, use `crossTanimotoSimilarityMemoryConstrained` (chunked compute, returns numpy on CPU).

### ETKDG conformer embedding

```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import ETKDGv3
from nvmolkit.embedMolecules import EmbedMolecules

mols = [AddHs(MolFromSmiles(smi)) for smi in ["C1CCCCC1", "C1CCCCC2CCCCC12", "COO"]]
params = ETKDGv3()
params.useRandomCoords = True

EmbedMolecules(mols, params, confsPerMolecule=10, maxIterations=-1)

for mol in mols:
    print(mol.GetNumConformers())
```

Inputs are `list[Mol]`, sanitized and with hydrogens added (`AddHs`). Conformers are added in-place. `params.useRandomCoords` must be `True` - nvMolKit's ETKDG only supports random-coord initialization. A handful of niche `EmbedParameters` options are not supported (bounds matrices, custom CPCI, coord maps, separate-fragment embedding); the Features section of the docs site lists the full restrictions.

### MMFF94 minimization of a batch of conformers

```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import ETKDGv3
from nvmolkit.embedMolecules import EmbedMolecules
from nvmolkit.mmffOptimization import MMFFOptimizeMoleculesConfs

mols = [AddHs(MolFromSmiles(smi)) for smi in ["CCO", "CCN", "c1ccccc1"]]
params = ETKDGv3(); params.useRandomCoords = True
EmbedMolecules(mols, params, confsPerMolecule=5)

energies = MMFFOptimizeMoleculesConfs(mols, maxIters=500)
for mol, mol_energies in zip(mols, energies):
    print(mol.GetNumConformers(), mol_energies)
```

Inputs are `list[Mol]` with conformers already populated (typically by ETKDG, RDKit's `EmbedMultipleConfs`, or a prior nvMolKit call). Coordinates are updated in place; the return is `list[list[float]]` of optimized energies aligned with the input molecule order and conformer index. UFF is identical in shape: swap in `from nvmolkit.uffOptimization import UFFOptimizeMoleculesConfs`.

If any input molecule is `None` or lacks MMFF/UFF atom types, the call raises `ValueError`. The exception's `args[1]` is a dict with keys `"none"` and `"no_params"` listing the offending indices - useful for filtering a noisy input set.

### Conformer RMSD and Butina clustering

```python
import torch
from rdkit import Chem
from rdkit.Chem.rdDistGeom import EmbedMultipleConfs
from nvmolkit.clustering import butina
from nvmolkit.conformerRmsd import GetConformerRMSMatrixBatch

mols = [Chem.AddHs(Chem.MolFromSmiles(smi)) for smi in ["CCCCCC", "c1ccccc1"]]
for mol in mols:
    EmbedMultipleConfs(mol, numConfs=10)

# Remove hydrogens after embedding for heavy-atom RMSD.
heavy_mols = [Chem.RemoveHs(mol) for mol in mols]

# Default RMSD output is RDKit-compatible condensed lower-triangle form.
condensed = GetConformerRMSMatrixBatch(heavy_mols)

# Butina expects a square distance matrix, so request square GPU tensors.
square = GetConformerRMSMatrixBatch(heavy_mols, output_format="square")
clusters = [butina(distance_matrix, cutoff=0.5).torch() for distance_matrix in square]

torch.cuda.synchronize()
for mol_clusters in clusters:
    print(mol_clusters.cpu().tolist())
```

`GetConformerRMSMatrix(mol)` and `GetConformerRMSMatrixBatch(mols)` default to `output_format="condensed"`, returning `AsyncGpuResult` objects that wrap RDKit-style flat vectors of length `N * (N - 1) // 2`. Use `output_format="square"` when chaining into `butina()` or any other API that expects an `N x N` distance matrix. Both forms live on the GPU; call `.numpy()` on condensed results or synchronize before moving square tensors to the CPU.

### Custom forcefield options + constraints (`BatchedForcefield`)

Reach for `MMFFBatchedForcefield` / `UFFBatchedForcefield` instead of the one-shot `MMFFOptimizeMoleculesConfs` / `UFFOptimizeMoleculesConfs` when you need any of:

- Custom `maxIters` / `forceTol` per call
- Per-molecule `nonBondedThreshold` (MMFF) or `vdwThreshold` (UFF), or per-molecule `ignoreInterfragInteractions`
- Per-molecule `MMFFMolProperties` objects (e.g. MMFF94s vs MMFF94)
- Distance, position, angle, or torsion constraints
- Standalone `compute_energy()` / `compute_gradients()` without minimization

```python
from rdkit.Chem import AddHs, MolFromSmiles
from rdkit.Chem.rdDistGeom import EmbedMultipleConfs
from nvmolkit.batchedForcefield import MMFFBatchedForcefield

mols = [AddHs(MolFromSmiles(smi)) for smi in ["CCO", "CCCCCC"]]
for mol in mols:
    EmbedMultipleConfs(mol, numConfs=5)

ff = MMFFBatchedForcefield(
    mols,
    nonBondedThreshold=[100.0, 20.0],
    ignoreInterfragInteractions=True,
)

ff[0].add_position_constraint(0, max_displ=0.1, force_constant=50.0)
ff[1].add_distance_constraint(0, 4, relative=False, min_len=1.8, max_len=2.2, force_constant=25.0)

energies, converged = ff.minimize(maxIters=500, forceTol=1e-4)
for mol, mol_energies, mol_converged in zip(mols, energies, converged):
    print(mol.GetNumConformers(), mol_energies, mol_converged)
```

All conformers of each input molecule are minimized in one batch. Constraints attached via `ff[i].add_*_constraint(...)` apply to every conformer of molecule `i`; constraint setters mark the wrapper dirty and the native forcefield rebuilds on the next call. Pass `output=CoordinateOutput.DEVICE` to `.minimize(...)` to keep optimized coordinates on the GPU (`Device3DResult`) instead of writing them back into RDKit conformers. UFF is the same shape: `UFFBatchedForcefield(mols, vdwThreshold=..., ...)`.

## Going deeper

- Full feature list, API reference, and guides: <https://nvidia-bionemo.github.io/nvMolKit/>
- What changed in each release: <https://nvidia-bionemo.github.io/nvMolKit/changelog.html>
- Worked examples (Jupyter notebooks): the `examples/` directory in the GitHub repo
openfold2-nim7.49 KB

View saved version →

---
name: openfold2-nim
description: >
  Use this skill for OpenFold2, NVIDIA's BioNeMo NIM microservice for monomer protein structure prediction. Invoke whenever the user mentions OpenFold2, AlphaFold2-like monomer folding, protein sequence-to-structure prediction, A3M MSAs, mmCIF templates, hosted NVIDIA API calls, or local Docker deployment.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# OpenFold2 NIM

Predict a single protein-chain structure from an amino-acid sequence, with
optional A3M multiple sequence alignments and mmCIF templates. Use this
`SKILL.md` for basic hosted/local NIM use; load supplemental files only when
the task needs deeper context:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: model scope, strengths, limitations, and handoffs.
- `references/parameters.md`: MSA, template, model-selection, and relax effects.
- `references/validation.md`: artifact and scientific sanity checks.
- `references/examples.md`: compact hosted/local payload patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted URL: `https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template`
- Local URL: `http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template`
- Local readiness: `http://localhost:8000/v1/health/ready`

Mode difference: hosted and local use the same prediction path except local
does not include `/v1/`. Hosted requests use `Authorization: Bearer
$NGC_API_KEY`; local inference requests use no auth header after readiness.

## Auth And Environment

Do not print API keys. Confirm they exist with shell tests, not echoes.

Hosted needs `NGC_API_KEY` in the request header. Supported local Docker
startup uses `NGC_API_KEY`, or `NVIDIA_API_KEY` as a fallback, plus
`LOCAL_NIM_CACHE`. A repo-root `.env` file may be sourced as a local override.

## Local Docker

Use the official OpenFold2 NIM image and mount `LOCAL_NIM_CACHE` at
`/opt/nim/.cache`. Current docs recommend at least 80 GB disk, 64 GB system
RAM, 8 CPU cores, and one supported GPU; the container is roughly 55 GB and
first startup downloads about 10 GB of model parameters.

When writing local setup commands, copy the preflight below exactly. Do not
drop `.env`, `NVIDIA_API_KEY`, `LOCAL_NIM_CACHE`, or the no-auth local request.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm --name openfold2 \
  --runtime=nvidia \
  --gpus "device=${NIM_TEST_GPU}" \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/openfold/openfold2:latest
```

Readiness check:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## Request Pattern

Use Python `requests`; curl escaping is fragile for A3M/mmCIF text. The
`sequence` field is required. `input_id`, `alignments`, `selected_models`,
`relax_prediction`, `use_templates`, and `explicit_templates` are optional.

```python
import os
import requests

hosted = True
url = (
    "https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template"
    if hosted
    else "http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template"
)
headers = {"Content-Type": "application/json"}
if hosted:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

seq = "MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT"
payload = {
    "sequence": seq,
    "input_id": "kras_fragment",
    "selected_models": [1],
    "relax_prediction": False,
    "alignments": {
        "uniref90": {
            "a3m": {
                "alignment": f">query\n{seq}",
                "format": "a3m",
            }
        }
    },
}

response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

Payload gotchas:

- OpenFold2 is monomer-only. For protein-ligand, protein-DNA/RNA, or
  multi-chain complexes, use OpenFold3 or Boltz2 instead.
- `sequence` must use valid amino-acid IUPAC symbols.
- Hosted API docs list sequence length 1-1000; local docs say current NIM
  supports sequences up to 2048 residues on supported hardware.
- A3M alignments go under `alignments` by database name, then `a3m` with
  `alignment` and `format`. When the user needs to create or deepen an MSA,
  hand off to `msa-search-nim` / MSA Search and map its A3M output into this
  `alignments` shape.
- Starting with OpenFold2 2.0.0, use `explicit_templates` with mmCIF content;
  do not write new HHR-template examples.
- `selected_models` chooses AlphaFold2/OpenFold parameter sets 1-5. Select one
  or two models for smoke tests; use all five for stronger production runs.

## Save And Interpret Output

The response includes one prediction per selected model, ordered by confidence.
Save every returned structure-like text field and the full JSON response so
field-shape differences are auditable. Production answers should explicitly
write `.pdb` or `.cif` artifacts, preserve the response JSON, and print any
confidence/ranking fields the service returns.

```python
from pathlib import Path
import json

Path("openfold2_response.json").write_text(json.dumps(result, indent=2))

def save_strings(obj, prefix="openfold2"):
    i = 0
    if isinstance(obj, dict):
        for key, value in obj.items():
            if isinstance(value, str) and ("ATOM" in value or value.lstrip().startswith("data_")):
                i += 1
                ext = "cif" if value.lstrip().startswith("data_") else "pdb"
                Path(f"{prefix}_{key}_{i}.{ext}").write_text(value)
            elif isinstance(value, (dict, list)):
                i += save_strings(value, f"{prefix}_{key}")
    elif isinstance(obj, list):
        for idx, value in enumerate(obj, start=1):
            if isinstance(value, (dict, list)):
                i += save_strings(value, f"{prefix}_{idx}")
    return i

saved = save_strings(result)
print(f"saved {saved} structure artifact(s)")
```

For production monomer runs:

- Use `selected_models: [1, 2, 3, 4, 5]` unless the user requests a smoke test.
- Use `relax_prediction: True` in Python payloads when relaxation is desired;
  JSON examples may show `true`.
- State the sequence length caveat: hosted API docs list 1-1000 residues, while
  local support-matrix docs list up to 2048 residues on supported hardware.
- If the task is a complex rather than a monomer, redirect to OpenFold3 or
  Boltz2.

Treat tiny toy sequences and single-sequence MSAs as API smoke tests, not
quality evidence. For scientific interpretation and validation, read
`references/science.md` and `references/validation.md`.

## Troubleshooting

- `401`: missing, expired, or unauthorized NGC API key.
- `422`: invalid amino-acid characters, sequence too long, malformed A3M, bad
  `selected_models`, or malformed mmCIF template object.
- Local `404`: remove `/v1/` from the prediction URL.
- Weak structures: use MSA Search to generate deeper A3M alignments and add
  biologically relevant mmCIF templates when appropriate.
- Local startup stalls: first run downloads parameters into `LOCAL_NIM_CACHE`.

Referenced files: 5

openfold3-nim7.08 KB

View saved version →

---
name: openfold3-nim
description: >
  Use this skill for OpenFold3, NVIDIA's BioNeMo NIM microservice for biomolecular structure prediction. Invoke whenever the user mentions OpenFold3 or needs protein, protein-ligand, protein-DNA/RNA, or multi-chain complex prediction with the hosted NVIDIA API or local Docker NIM. Covers endpoint choice, auth, request payloads, output artifacts, confidence scores, and local container setup.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# OpenFold3 NIM

Predict biomolecular structures with OpenFold3. It supports proteins, DNA, RNA,
small-molecule ligands, and multi-entity assemblies. Use this `SKILL.md` for
basic hosted/local NIM use; load supplemental files only when the task needs
deeper context:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: purpose, strengths, limitations, and model handoffs.
- `references/parameters.md`: molecule fields, MSAs, templates, samples, tuning.
- `references/validation.md`: artifact checks and scientific sanity checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted URL: `https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict`
- Local URL: `http://localhost:8000/biology/openfold/openfold3/predict`
- Local readiness: `http://localhost:8000/v1/health/ready`

Mode difference: the local prediction path has no `/v1/` prefix. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Auth And Environment

Do not print API keys. Confirm they exist with shell tests, not echoes.

Hosted needs `NGC_API_KEY` in the request header. Local startup needs
`NGC_API_KEY`, or `NVIDIA_API_KEY` as a fallback, plus `LOCAL_NIM_CACHE`.
A repo-root `.env` file may be sourced as a local override before validation.

## Local Docker

Use the official OpenFold3 NIM image and mount `LOCAL_NIM_CACHE` at
`/opt/nim/.cache`. First startup downloads model artifacts and can take several
minutes.

When writing local setup commands, copy the preflight below exactly. Do not
replace it with a simple `: "${NGC_API_KEY:?Set NGC_API_KEY}"` check, do not
drop `NVIDIA_API_KEY`, and do not invent a default `LOCAL_NIM_CACHE`; those
lines are the repo's local NIM env contract. The default single-GPU launch
should show the literal `--gpus "device=0"`; choose a different device only
when the user asks.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run --rm --name openfold3 \
  --runtime=nvidia \
  --gpus "device=0" \
  --shm-size=16g \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/openfold/openfold3:latest
```

Readiness check:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## Request Pattern

Use `requests.post(..., json=payload, timeout=300)`. For local Docker tasks,
set `hosted = False` after the readiness check passes.

```python
import os
import requests

hosted = True
url = (
    "https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict"
    if hosted
    else "http://localhost:8000/biology/openfold/openfold3/predict"
)
headers = {"Content-Type": "application/json"}
if hosted:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

seq = "MKTVRQERLKSIVR"
payload = {
    "inputs": [{
        "input_id": "prediction_1",
        "output_format": "pdb",
        "molecules": [{
            "type": "protein",
            "id": "A",
            "sequence": seq,
            "diffusion_samples": 1,
            "msa": {
                "main": {
                    "a3m": {
                        "alignment": f">query\n{seq}",
                        "format": "a3m"
                    }
                }
            }
        }]
    }]
}

response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

Payload gotchas:

- Top level is `{"inputs": [...]}` and OpenFold3 accepts exactly one input.
- `molecules` can contain 1-32 objects with `type`: `protein`, `dna`, `rna`,
  or `ligand`.
- Protein/RNA MSAs are optional but, when supplied, `alignment` must start with
  a FASTA header such as `>query\nSEQUENCE`.
- Ligands use either `smiles` or `ccd_codes`, for example
  `{"type": "ligand", "id": "L", "ccd_codes": "ATP"}`.
- DNA/RNA entities use `sequence`, for example
  `{"type": "dna", "id": "B", "sequence": "ATCGATCG"}`.
- `diffusion_samples` is 1-5. `output_format` is `pdb` or `cif`.

## Save And Interpret Output

Save every returned structure as a scientific artifact. Main response path:
`result["outputs"][0]["structures_with_scores"]`.

```python
output = result["outputs"][0]
for i, sample in enumerate(output["structures_with_scores"], start=1):
    fmt = sample["format"]
    with open(f"openfold3_structure_{i}.{fmt}", "w", encoding="utf-8") as fh:
        fh.write(sample["structure"])
    print("confidence_score", sample.get("confidence_score"))
    print("complex_plddt_score", sample.get("complex_plddt_score"))
    print("ptm_score", sample.get("ptm_score"))
    print("iptm_score", sample.get("iptm_score"))
    print("complex_pde_score", sample.get("complex_pde_score"))
```

Higher `confidence_score`, `complex_plddt_score`, `ptm_score`, and `iptm_score`
are generally better; lower `complex_pde_score` is generally better. Treat toy
or very short sequences as API smoke tests, not meaningful structural biology.
For why and when OpenFold3 is scientifically appropriate, read
`references/science.md`.

## Common Limits

- Inputs per request: 1.
- Molecules per input: 1-32.
- Diffusion samples: 1-5.
- TensorRT path supports shorter sequences; PyTorch path can support longer
  sequences, but long inputs need much more GPU memory.
- Sequences over roughly 1800 residues require at least 80 GB GPU memory.
- Local NIM is single-GPU only; choose the target device in the Docker flag.

## Troubleshooting

- `401`: missing, expired, or unauthorized NGC API key.
- `422`: invalid molecule type, invalid sequence characters, bad MSA shape, or
  `diffusion_samples` outside 1-5.
- MSA errors: ensure the alignment starts with `>query\n`.
- Local `404`: remove `/v1/` from the prediction URL.
- Local startup stalls: first run may be downloading 10-15 GB of model weights
  into `LOCAL_NIM_CACHE`.
- Memory errors: shorten the sequence, reduce samples, or use a larger GPU.

Referenced files: 5

parabricks8.19 KB

View saved version →

---
name: parabricks
description: >-
  Route NVIDIA Parabricks pbrun tools, assess GPU/runtime readiness, and provide
  version-aware command guidance for FASTQ/BAM processing, RNA-seq, variant
  calling, BAM QC, and GVCF workflows. Do NOT use for inspecting or accelerating
  whole pipelines — use genomics-workflow-acceleration.
license: CC-BY-4.0 AND Apache-2.0
metadata:
  version: "1.1.0"
  tags:
    - parabricks
    - genomics
    - nvidia
---

# Parabricks

## Purpose

Use this skill to discover the right NVIDIA Parabricks `pbrun` command, assess
runtime readiness, and generate version-aware command guidance for individual
tools and pipelines.

Do **not** use this skill for whole-workflow inspection, acceleration planning,
or wiring optional GPU branches. For pipeline-level work, use
`genomics-workflow-acceleration`.

## When to Use This Skill

- Which `pbrun` tool fits the user's data and goal
- GPU, driver, Docker, container, storage, or installation readiness
- Command shape, flags, and validation for a specific Parabricks tool
- Troubleshooting a single Parabricks command or tool family

## Prerequisites

Ask for input data type, sequencing technology, reference build, sample
structure, desired output, target Parabricks version/container tag, and runtime
target before recommending commands.

If the user is unsure which tool applies, read
[tool-index.md](references/tool-index.md) first, then load the matching
`references/pbrun-<tool>.md` file.

## Limitations

This skill routes and guides Parabricks commands. It does not install
Parabricks, infer missing sample metadata, guarantee output parity, provide
clinical interpretation, or promise exact runtime without benchmark data.

## Workflow

1. Confirm the Parabricks version or container tag. Verify the current NVIDIA
   docs when the user asks for the latest tool list or version-sensitive flags.
2. Classify the request:
   - **Runtime** → [runtime-environment.md](references/runtime-environment.md)
   - **Tool discovery** → [tool-index.md](references/tool-index.md)
   - **Specific command** → matching `references/pbrun-<tool>.md`
3. Collect missing biological and filesystem context before generating commands.
4. Generate conservative Docker commands with explicit mounts, workdir, and
   placeholders. Validate paths, indexes, and outputs after command generation.

## Tool Reference Index

Load only the reference file for the selected tool.

| Tool | Reference | Use when |
|------|-----------|----------|
| `applybqsr` | [pbrun-applybqsr.md](references/pbrun-applybqsr.md) | Apply BQSR table to aligned BAM |
| `bam2fq` | [pbrun-bam2fq.md](references/pbrun-bam2fq.md) | BAM → FASTQ conversion |
| `bamsort` | [pbrun-bamsort.md](references/pbrun-bamsort.md) | Standalone BAM sort |
| `bqsr` | [pbrun-bqsr.md](references/pbrun-bqsr.md) | Generate BQSR recalibration table |
| `fq2bam` | [pbrun-fq2bam.md](references/pbrun-fq2bam.md) | Short-read DNA paired FASTQ → BAM/CRAM |
| `fq2bam_meth` | [pbrun-fq2bam_meth.md](references/pbrun-fq2bam_meth.md) | Bisulfite/methylation FASTQ → BAM/CRAM |
| `giraffe` | [pbrun-giraffe.md](references/pbrun-giraffe.md) | Pangenome graph alignment |
| `markdup` | [pbrun-markdup.md](references/pbrun-markdup.md) | Standalone duplicate marking |
| `minimap2` | [pbrun-minimap2.md](references/pbrun-minimap2.md) | Long-read FASTQ alignment |
| `rna_fq2bam` | [pbrun-rna_fq2bam.md](references/pbrun-rna_fq2bam.md) | RNA-seq FASTQ(s) → splice-aware BAM (STAR alignment) |
| `starfusion` | [pbrun-starfusion.md](references/pbrun-starfusion.md) | Fusion detection from chimeric junction input + STAR-Fusion genome library |
| `germline` | [pbrun-germline.md](references/pbrun-germline.md) | GATK-style germline pipeline from FASTQ |
| `deepvariant_germline` | [pbrun-deepvariant_germline.md](references/pbrun-deepvariant_germline.md) | DeepVariant germline pipeline from FASTQ |
| `haplotypecaller` | [pbrun-haplotypecaller.md](references/pbrun-haplotypecaller.md) | Standalone HaplotypeCaller from BAM/CRAM |
| `deepvariant` | [pbrun-deepvariant.md](references/pbrun-deepvariant.md) | Standalone DeepVariant from BAM/CRAM |
| `somatic` | [pbrun-somatic.md](references/pbrun-somatic.md) | Tumor-normal somatic pipeline |
| `mutectcaller` | [pbrun-mutectcaller.md](references/pbrun-mutectcaller.md) | Mutect2-compatible somatic calling |
| `deepsomatic` | [pbrun-deepsomatic.md](references/pbrun-deepsomatic.md) | DeepSomatic-based somatic calling |
| `pacbio_germline` | [pbrun-pacbio_germline.md](references/pbrun-pacbio_germline.md) | PacBio long-read germline |
| `ont_germline` | [pbrun-ont_germline.md](references/pbrun-ont_germline.md) | Oxford Nanopore long-read germline |
| `pangenome_germline` | [pbrun-pangenome_germline.md](references/pbrun-pangenome_germline.md) | Pangenome-aware germline |
| `pangenome_aware_deepvariant` | [pbrun-pangenome_aware_deepvariant.md](references/pbrun-pangenome_aware_deepvariant.md) | Pangenome-aware DeepVariant |
| `prepon` | [pbrun-prepon.md](references/pbrun-prepon.md) | Pangenome-aware preprocessing |
| `postpon` | [pbrun-postpon.md](references/pbrun-postpon.md) | Pangenome-aware post-processing |
| `bammetrics` | [pbrun-bammetrics.md](references/pbrun-bammetrics.md) | Whole-genome coverage/depth metrics |
| `collectmultiplemetrics` | [pbrun-collectmultiplemetrics.md](references/pbrun-collectmultiplemetrics.md) | Multiple Picard/GATK-style alignment metrics |
| `genotypegvcf` | [pbrun-genotypegvcf.md](references/pbrun-genotypegvcf.md) | Joint-genotype GVCF input(s) into VCF |
| `indexgvcf` | [pbrun-indexgvcf.md](references/pbrun-indexgvcf.md) | Index GVCF input |
| `dbsnp` | [pbrun-dbsnp.md](references/pbrun-dbsnp.md) | dbSNP annotation on variant files |

For routing heuristics when multiple tools could apply, see
[tool-index.md](references/tool-index.md).

## Runtime Readiness

For GPU, driver, Docker, container, storage, or installation questions, read
[runtime-environment.md](references/runtime-environment.md) and prefer:

```bash
python3 skills/parabricks/scripts/check_parabricks_runtime.py
```

Add `--path <dir>` for known input/output/tmp paths. Run container probes only
with user consent.

## Command Shape

```bash
docker run --rm --gpus all \
  --volume /host/input:/workdir \
  --volume /host/output:/outputdir \
  --workdir /workdir \
  nvcr.io/nvidia/clara/clara-parabricks:<version> \
  pbrun <selected-tool> \
  <tool-specific-options>
```

Check the version-specific tool reference before finalizing flags.

## Troubleshooting

| Error | Cause | Solution |
|-------|-------|----------|
| Multiple plausible tools | Data type or goal underspecified | Ask for assay, inputs, caller preference, desired output; use tool-index |
| Exact flag requested | Options are version-sensitive | Check the selected tool reference and NVIDIA docs |
| Runtime question | GPU, Docker, drivers, or storage | Use runtime-environment reference and diagnostic script |
| Wrong tool family | Assay or input type unclear | Confirm DNA/RNA/methylation/long-read/pangenome before routing |
| CUDA or memory failure | Runtime not ready or GPU memory constrained | Assess runtime before tuning command flags |

## Guardrails

- Treat command availability and options as version-sensitive.
- Do not infer exact flags from command names alone.
- Do not collapse standalone tools and full pipelines when explaining tradeoffs.
- Do not substitute DNA `fq2bam` for RNA, or germline for somatic callers.
- Do not invent sample names, read groups, reference builds, known-sites files,
  model files, graph resources, container tags, or output paths.
- Do not install, upgrade, or modify packages. Label setup commands as user-run.
- Do not claim CPU execution of Parabricks tools.
- Do not claim biological or VCF parity without a comparison run.
- Prefer official NVIDIA docs for exact command syntax and option defaults.

## Key References

- Parabricks tool index:
  <https://docs.nvidia.com/clara/parabricks/latest/toolreference.html>
- Output accuracy and compatible CPU software versions:
  <https://docs.nvidia.com/clara/parabricks/latest/documentation/tooldocs/outputaccuracyandcompatiblecpusoftwareversions.html>
- Getting started:
  <https://docs.nvidia.com/clara/parabricks/latest/gettingstarted.html>
- Overview:
  <https://docs.nvidia.com/clara/parabricks/latest/overview.html>

Referenced files: 33

protein-binder-design5.66 KB

View saved version →

---
name: protein-binder-design
description: >
  Orchestrate an end-to-end de novo protein binder design campaign against a protein target by composing BioNeMo NIM skills. Use for binder design, minibinder design, de novo binders, RFdiffusion + ProteinMPNN + Boltz2/OpenFold3 pipelines, epitope/hotspot-targeted design, in-silico binder validation, and ranking designs by interface confidence.
license: Apache-2.0
compatibility: "numpy>=1.24; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# Protein Binder Design (workflow)

Run a de novo binder design campaign by composing atomic NIM skills. This skill
owns orchestration, handoff contracts, filtering, validation, and the run
manifest. It does NOT duplicate per-NIM API details — defer those to each
atomic skill's `SKILL.md`.

## Composed skills

| Step | Skill | Owns |
|---|---|---|
| Backbones | `rfdiffusion-nim` | binder backbone PDBs (contigs + hotspots) |
| Sequences | `proteinmpnn-nim` | sequences for each backbone |
| Co-fold / score | `boltz2-nim` or `openfold3-nim` | binder–target complex + confidence / ipTM |
| MSA (optional) | `msa-search-nim` | target A3M for higher-quality folding |

The atomic NIM skills are recommended companions (one per NIM, from the BioNeMo
NIM skill set). They are **not required**: `references/pipeline.md` carries the
concrete request shape for every NIM call, so an agent with NIM access can follow
this skill standalone. For endpoints/auth see **Configuration** below.

## Pipeline

1. **Target prep** — get the target PDB + epitope; map epitope/hotspot author
   residue numbers to RFdiffusion `hotspot_res` strings; optionally build a
   target MSA with `msa-search-nim`.
2. **Backbones** (`rfdiffusion-nim`) — binder contig + `hotspot_res`; N backbones.
3. **Sequences** (`proteinmpnn-nim`) — k sequences per backbone; drop the
   native/WT row from `mfasta`.
4. **Co-fold + score** (`boltz2-nim` / `openfold3-nim`) — co-fold binder+target;
   collect interface confidence (ipTM) and binder pLDDT.
5. **Self-consistency** — CA-RMSD between the RFdiffusion backbone and the
   predicted binder (`scripts/metrics.py`).
6. **Filter + rank** — apply thresholds; rank survivors; write manifest + CSV.

Full handoff contracts, branching, and the cost funnel: `references/pipeline.md`.

## Handoff contracts (the fragile glue)

- RFdiffusion `output_pdb` → ProteinMPNN `input_pdb` (inline PDB text).
- ProteinMPNN `mfasta` → Boltz2 binder polymer `sequence` (exclude the
  native/WT row; pair scores only with designed rows).
- Epitope author residue numbers → 1-based sequence indices: remap with
  `scripts/pdb_utils.py:remap_to_seq_index`. RFdiffusion `hotspot_res` uses
  chain+author strings like `"A50"`; Boltz2 pocket/contacts use 1-based indices.
- Boltz2 complex `.cif` → binder chain → self-consistency RMSD vs the backbone.

## Run manifest (reproducibility backbone)

Every campaign writes `manifest.json` (+ `candidates.csv`) under a run dir via
`scripts/manifest.py`. It records lineage, params, scores, artifacts, filter
status, and controls — enabling ranking, resumability, validation, and the
final report. Schema and usage: `references/manifest.md`.

## Filters (defaults)

- ipTM ≥ 0.8, binder pLDDT ≥ 80, self-consistency RMSD ≤ 2.0 Å.
- Override per campaign and record overrides in the manifest `filters`.

## Validation

Always run controls and report a **success rate**, not just top scores.
Negative controls via `scripts/controls.py` (scrambled sequences); positive
controls = published binders re-scored through the same pipeline. Benchmark
targets live in `assets/targets.json` (`scripts/registry.py`). Methodology and
metric definitions: `references/validation.md`.

## Human-in-the-loop + cost

- Confirm target, epitope/hotspots, binder length range, and hosted-vs-local
  with the user before generating backbones (AskUserQuestion).
- Co-folding is the expensive stage: co-fold a capped shortlist, review, then
  expand. State hosted vs local once and reuse it across all NIM calls.

## Responsible use

De novo binder design is dual-use. Decline requests aimed at enhancing pathogen
fitness, toxin potency, or bioweapon function; keep designs to legitimate
research and therapeutic intent.

## Configuration (NIM access)

Each composed NIM is reached over HTTP; choose **hosted** or **local** once and
reuse it for every call:

- **Hosted** (managed): base URL `https://health.api.nvidia.com/v1/...` per NIM at
  [build.nvidia.com](https://build.nvidia.com); set `NVIDIA_API_KEY` (sent as
  `Authorization: Bearer`). Read keys from the env — never hardcode them.
- **Local** (self-hosted NGC containers): point each NIM at its local URL
  (e.g. `http://localhost:8000/...`); local NIMs need no auth header. To **launch** the
  NIMs yourself (docker run per NIM, persistent caches, health checks, and the GPU
  **profile‑selection gotcha** — some NIMs (e.g. Boltz2) need `NIM_MODEL_PROFILE` pinned
  on GPUs that have no bundled profile, while others (RFdiffusion/ProteinMPNN) auto‑select
  by compute capability): see **`references/local-nim-setup.md`**.

Per-NIM paths, request/response schemas, and worked `curl`/Python examples live in
`references/pipeline.md`.

## Scripts

- `scripts/manifest.py` — campaign manifest (create / load / score / filter / rank / CSV).
- `scripts/pdb_utils.py` — PDB parse, chain extract, sequence, residue remap, CA coords.
- `scripts/metrics.py` — Kabsch CA-RMSD for self-consistency.
- `scripts/controls.py` — scrambled negative controls.
- `scripts/registry.py` + `assets/targets.json` — **example** benchmark target
  registry (illustrative epitopes — verify against the cited structure before a
  real campaign). Replace with your own targets.

Referenced files: 12

proteinmpnn-nim4.78 KB

View saved version →

---
name: proteinmpnn-nim
description: >
  Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Use for ProteinMPNN, inverse folding, sequence design, backbone redesign, fixed chains/residues, omit_AAs, sampling temperature, soluble model, hosted NVIDIA API, local Docker, PDB input, and multi-FASTA output.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# ProteinMPNN NIM

Design protein sequences for a supplied backbone PDB. Use this `SKILL.md` for
first-pass hosted/local usage; load supplemental files only when needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: inverse-folding uses, limits, and validation.
- `references/parameters.md`: design controls, fixed positions, sampling.
- `references/validation.md`: FASTA, score, and structure checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`
- Local: `http://localhost:8000/biology/ipd/proteinmpnn/predict`

Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker

For local setup answers, copy the preflight below exactly before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not answer with only
a localhost Python request. This NIM's cache mount is unique:
`/home/nvs/.cache/nim`, not `/opt/nim/.cache`.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

export NIM_TEST_GPU="${NIM_TEST_GPU:-0}"
mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run -it \
  --runtime=nvidia \
  --gpus "device=${NIM_TEST_GPU}" \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/home/nvs/.cache/nim" \
  -p 8000:8000 \
  nvcr.io/nim/ipd/proteinmpnn:latest
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## Request Pattern

Read PDB content inline; do not send only a file path.

```python
import os
from pathlib import Path
import requests

HOSTED = True
pdb_content = Path("1R42.pdb").read_text()
url = (
    "https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
    if HOSTED else "http://localhost:8000/biology/ipd/proteinmpnn/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "input_pdb": pdb_content,
    "num_seq_per_target": 10,
    "sampling_temp": [0.1],
    "use_soluble_model": False,
    "ca_only": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```

Common controls:

- Redesign only chain A: `"input_pdb_chains": ["A"]`.
- Exclude amino acids: `"omit_AAs": ["C"]` or `"omit_AAs": ["M"]`.
- Diversity: `"sampling_temp": [0.1, 0.3, 0.5]` (always a list).
- Solubility bias: `"use_soluble_model": True`.
- Candidate count: `num_seq_per_target` is 1-100.

## Save And Report Output

```python
mfasta = result["mfasta"]
Path("designed_sequences.fa").write_text(mfasta)

# Scores correspond to designed sequences. The mfasta may include a native/WT
# row; do not pair that row with generated-sequence scores.
headers = [line for line in mfasta.splitlines() if line.startswith(">")]
designed_headers = [
    h for h in headers if "native" not in h.lower() and "wt" not in h.lower()
]
for header, score in zip(designed_headers, result.get("scores", [])):
    print(f"{header} score: {score:.4f}")
```

Validate promising designs by predicting structures with Boltz2 or OpenFold3
and comparing them to the target backbone. For FASTA/score sanity checks, read
`references/validation.md`.

## Limits And Troubleshooting

- Minimum GPU VRAM: about 3 GB.
- `sampling_temp` must be a list, even for one value.
- Empty `mfasta`: check non-empty `input_pdb` and `num_seq_per_target >= 1`.
- PDB parse errors: use valid PDB ATOM records.
- Local URL 404 usually means an accidental `/v1/` prefix.
- Cache mount error: use `/home/nvs/.cache/nim` inside the container.

Referenced files: 5

rfdiffusion-nim5.3 KB

View saved version →

---
name: rfdiffusion-nim
description: >
  Run RFDiffusion protein backbone design via NVIDIA NIM. Use for de novo protein backbones, motif scaffolding, binder design, hotspot residues, contigs syntax, diffusion steps, hosted NVIDIA API calls, local Docker deployment, and PDB backbone outputs for ProteinMPNN sequence design.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
---

# RFDiffusion NIM

Design protein backbone PDBs for de novo proteins, motif scaffolds, and binders.
Use this `SKILL.md` for first-pass hosted/local usage; load supplemental files
only when needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
- `references/science.md`: design modes, strengths, limits, and handoffs.
- `references/parameters.md`: contigs, hotspots, steps, and seeds.
- `references/validation.md`: PDB, contig, and artifact sanity checks.
- `references/examples.md`: compact hosted/local request patterns.

## Choose Mode

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/rfdiffusion/generate`
- Local: `http://localhost:8000/biology/ipd/rfdiffusion/generate`

Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
registry login, entitlement checks, and first-run model downloads; pass it
into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Local Docker

For local setup answers, copy the preflight below exactly before `docker login`,
`docker run`, readiness, and the no-auth local request. Do not replace it with a
simple `: "${NGC_API_KEY:?Set NGC_API_KEY}"` check, do not invent a cache
default, and do not drop the `NVIDIA_API_KEY` fallback. Default setup is single
GPU `device=0`.

```bash
set -a
[ -f .env ] && . ./.env
set +a

if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

mkdir -p "${LOCAL_NIM_CACHE}"
chmod 777 "${LOCAL_NIM_CACHE}"

docker run -it \
  --runtime=nvidia \
  --gpus "device=0" \
  -e NGC_API_KEY \
  -v "${LOCAL_NIM_CACHE}:/opt/nim/.cache" \
  -p 8000:8000 \
  nvcr.io/nim/ipd/rfdiffusion:2
```

Readiness:

```bash
until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done
```

## Contigs DSL

`contigs` defines what to keep and what to generate.

- `"100"`: generate exactly 100 residues.
- `"80-120"`: generate 80-120 residues.
- `"A25-35"`: keep chain A residues 25-35 from `input_pdb`.
- `"A25-35/0 50-80"`: keep A25-35, insert chain break `/0`, generate 50-80.

Design modes:

- De novo: `contigs="80-120"`; live hosted validation requires a non-empty
  `input_pdb` or `input_pdb_asset`, so inline requests should include the dummy
  PDB below.
- Motif scaffolding: read `target.pdb`, pass `input_pdb`, use a contig like
  `"A25-35/0 50-80"`.
- Binder design: pass target `input_pdb`, contig with target and binder segment,
  and `hotspot_res=["A50", "A51", ...]` in ChainResidue string format.

```python
DUMMY_PDB = (
    "CRYST1    1.000    1.000    1.000  90.00  90.00  90.00 P 1           1\n"
    "ATOM      1  CA  ALA A   1       0.000   0.000   0.000  1.00  0.00           C\n"
    "END\n"
)
```

## Request Pattern

```python
import os
from pathlib import Path
import requests

HOSTED = True
url = (
    "https://health.api.nvidia.com/v1/biology/ipd/rfdiffusion/generate"
    if HOSTED else "http://localhost:8000/biology/ipd/rfdiffusion/generate"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
    headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
    "input_pdb": DUMMY_PDB,
    "contigs": "80-120",
    "diffusion_steps": 50,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
Path("designed_backbone.pdb").write_text(result["output_pdb"])
```

Motif scaffold:

```python
payload = {
    "input_pdb": Path("target.pdb").read_text(),
    "contigs": "A25-35/0 50-80",
    "diffusion_steps": 50,
}
```

Binder design:

```python
payload = {
    "input_pdb": Path("target.pdb").read_text(),
    "contigs": "A1-100/0 50-100",
    "hotspot_res": ["A50", "A51", "A52", "A53", "A54"],
    "diffusion_steps": 50,
}
```

## Save And Interpret Output

Save `result["output_pdb"]` as a PDB artifact and report `elapsed_ms` when
present. Generated backbones are not final proteins; feed them to ProteinMPNN
for sequence design, then validate sequences/structures with Boltz2 or
OpenFold3. For PDB and contig checks, read `references/validation.md`.

## Limits And Troubleshooting

- `diffusion_steps`: 1-50; 50 is maximum quality, fewer is faster.
- Single GPU; minimum GPU VRAM is about 12 GB.
- `hotspot_res` uses strings like `"A50"`, not tuples.
- `422` usually means chain IDs in `contigs`/`hotspot_res` do not match
  `input_pdb`, a malformed contig, or omitted `input_pdb` for hosted de novo.
- Local URL 404 usually means an accidental `/v1/` prefix.

Referenced files: 5

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 12:00 UTC
Collection status
Collected

plugins_6a76572d8f8081918362aa7ff90947fb

Download listing JSON