← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORY

Update to NVIDIA BioNeMo Agent Toolkit

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "complexa-design",
  "description": "End-to-end Proteina-Complexa design pipeline driver. Reach for this skill whenever the user wants to \"design a binder\", \"design binders for X\", \"run complexa design\", \"de novo binder\", \"PDL1 binder\", \"TrkA binder\", \"design proteins for target\", \"protein binder design\", \"ligand binder\", \"design a small-molecule binder\", \"ATP-binding protein\", \"AME motif scaffolding\", \"scaffold a motif near a ligand\", \"motif + ligand design\", \"enzyme scaffolding\", \"flow matching protein design\", \"beam-search binder\", \"FK steering\", \"MCTS protein design\", \"refold with AF2\", \"refold with RF3\", or wants success rates, interface pAE, scRMSD, or FoldSeek diversity from a single command. This is the scientific anchor of the skill set: it drives `complexa design <pipeline>` from target picking to manifest emission and tells the user how many designs passed.",
  "included_files": [
    {
      "relative_path": "reference/overrides.md",
      "size_in_bytes": 23684
    },
    {
      "relative_path": "reference/pipelines.md",
      "size_in_bytes": 8917
    },
    {
      "relative_path": "reference/troubleshooting.md",
      "size_in_bytes": 9507
    }
  ],
  "skill_md_contents": "---\nname: complexa-design\ndescription: >\n  End-to-end Proteina-Complexa design pipeline driver. Reach for this skill whenever\n  the user wants to \"design a binder\", \"design binders for X\", \"run complexa\n  design\", \"de novo binder\", \"PDL1 binder\", \"TrkA binder\", \"design proteins for\n  target\", \"protein binder design\", \"ligand binder\", \"design a small-molecule\n  binder\", \"ATP-binding protein\", \"AME motif scaffolding\", \"scaffold a motif\n  near a ligand\", \"motif + ligand design\", \"enzyme scaffolding\", \"flow matching\n  protein design\", \"beam-search binder\", \"FK steering\", \"MCTS protein design\",\n  \"refold with AF2\", \"refold with RF3\", or wants success rates, interface pAE,\n  scRMSD, or\n  FoldSeek diversity from a single command. This is the scientific anchor of\n  the skill set: it drives `complexa design <pipeline>` from target picking to\n  manifest emission and tells the user how many designs passed.\ncompatibility: \"complexa CLI installed (pip install -e .); .env populated; 1x CUDA GPU >=40GB VRAM (A100/H100/L40S); 24 CPUs; ~50GB disk\"\nallowed-tools: Bash, Read, Write, AskUserQuestion\n---\n\n# Complexa Design Skill\n\nDrive the full four-stage `complexa design` pipeline: generate (flow matching +\nsearch) -> filter (top-N by reward) -> evaluate (refold with AF2/RF3) ->\nanalyze (success rate, FoldSeek/MMseqs diversity). Pick the right pipeline\nconfig for the design intent, validate the run upfront so the user does not\ndiscover a missing ckpt mid-folding, run it, and emit a replayable manifest +\nper-design success CSV.\n\n## What this skill enables\n\n- Protein binder design for protein targets (AF2 reward + ColabDesign refold).\n- Ligand binder design for small-molecule targets (RF3 reward + RF3 refold).\n- AME motif scaffolding with ligand context (motif + ligand features, RF3).\n- Search-based optimization: single-pass, best-of-n, beam-search, fk-steering, mcts.\n- Refold backends: ColabDesign (AF2), RF3, Boltz2, ESMFold (fast iteration).\n- Pass-rate + diversity analysis with per-`result_type` thresholds.\n\n## Step 1: Pre-flight\n\nAlways run the shared preflight before launching a design — generation needs the\nGPU and the right checkpoint, evaluation needs AF2/RF3 weights and tool\nbinaries. Bail early if the host cannot run the chosen pipeline.\n\n```bash\nbash .claude/skills/_shared/scripts/preflight.sh\n```\n\nRead `./complexa_setup/preflight.json` and bail if any of these are missing for\nthe chosen pipeline:\n\n- `gpu.available: false` -> all pipelines fail.\n- `gpu.vram_gb < 40` -> generation OOMs at default `batch_size: 16`; lower to 8.\n- `ckpts.complexa[.ckpt]` -> required for protein binder.\n- `ckpts.complexa_ligand[.ckpt]` -> required for ligand binder.\n- `ckpts.complexa_ame[.ckpt]` -> required for AME.\n- `env.AF2_DIR` missing -> protein binder default eval (`colabdesign`) fails.\n- `env.RF3_CKPT_PATH` or `env.RF3_EXEC_PATH` missing -> ligand binder / AME default eval (`rf3_latest`) fails.\n\nIf a ckpt is missing, point at `complexa-setup` and have the user run\n`complexa download --complexa-<variant>` first.\n\n## Step 2: Pick the pipeline\n\nComplexa has **one default pipeline (protein binder) and two extensions**\n(ligand binder, AME / enzyme). The pipeline is selected entirely\nby the `configs/search_*_pipeline.yaml` you pass to `complexa design` — each\nYAML pins its own model checkpoint, autoencoder, targets dict, default reward,\nand default refold backend. Switching pipelines is just \"swap the config path\nand the target name comes from a different dict\".\n\n### Default — protein binder\n\n```bash\ncomplexa design configs/search_binder_local_pipeline.yaml \\\n    ++run_name=pdl1_v1 ++generation.task_name=02_PDL1\n```\n\nUse this when the user says \"design a binder for X\", \"PDL1 binder\",\n\"de novo binder\", \"design proteins for a target\", etc. — i.e. the target is\na protein surface. The config pins `complexa.ckpt` + `complexa_ae.ckpt`, reads\ntargets from `configs/targets/targets_dict.yaml`, rewards with AF2\n(`af2folding`), inverse-folds with SolubleMPNN, and evaluates with\nColabDesign / AF2. **If the user did not specify, this is what they want.**\n\n### Extension A — ligand binder (small-molecule pocket)\n\n```bash\ncomplexa design configs/search_ligand_binder_local_pipeline.yaml \\\n    ++run_name=v11_v1 ++generation.task_name=39_7V11_LIGAND \\\n    ++metric.binder_folding_method=rf3_latest\n```\n\nSwitch to this when the user says \"ligand binder\", \"small-molecule pocket\",\n\"SMILES target\", \"ATP-binding protein\", or names a target ending in\n`_LIGAND` / from the FAD / SAM / OQO / 7V11 / etc. families. The config\n**also activates LoRA** (`r=32`, `lora_alpha=64`) which is required for the\nreleased ligand checkpoint — leave the `lora:` block alone.\n\n### Extension B — AME / motif + ligand (enzyme scaffolding)\n\n```bash\ncomplexa design configs/search_ame_local_pipeline.yaml \\\n    ++run_name=ame_chm ++generation.task_name=M0096_1chm\n```\n\nSwitch to this when the user says \"scaffold a motif near a ligand\",\n\"active-site design\", \"enzyme scaffolding\", \"AME\", or names a target like\n`M0024_1nzy`, `M0096_1chm` (the `M####_<pdb>` AME task naming). The config sets\n`env_vars.USE_V2_COMPLEXA_ARCH=True` (the CLI runner injects it into the\nsubprocess) and uses both `MotifFeatures` and `LigandFeatures`. Default search\nis `single-pass`; switch to `best-of-n` only if you also enable the\n`CompositeRewardModel` (commented out in `ame_generate.yaml`).\n\n### Pipeline cheat sheet — what changes when you switch\n\n| Knob | Protein binder (default) | Ligand binder | AME (enzyme) |\n|---|---|---|---|\n| **Pipeline YAML** | `configs/search_binder_local_pipeline.yaml` | `configs/search_ligand_binder_local_pipeline.yaml` | `configs/search_ame_local_pipeline.yaml` |\n| **Model ckpt** | `complexa.ckpt` | `complexa_ligand.ckpt` | `complexa_ame.ckpt` |\n| **Autoencoder ckpt** | `complexa_ae.ckpt` | `complexa_ligand_ae.ckpt` | `complexa_ame_ae.ckpt` |\n| **Targets dict** | `configs/targets/targets_dict.yaml` | `configs/targets/ligand_targets_dict.yaml` | `configs/design_tasks/ame_dict_v2.yaml` |\n| **Task-name pattern** | `<NN>_<NAME>` (e.g. `02_PDL1`, `22_DerF21`) | `<NN>_<PDB>_LIGAND` (e.g. `39_7V11_LIGAND`) | `M####_<pdb>` (e.g. `M0096_1chm`) |\n| **`USE_V2_COMPLEXA_ARCH`** | (unset → v1) | (unset → v1) | `\"True\"` (set in YAML) |\n| **LoRA** | (none) | required (`r=32, alpha=64`) | required |\n| **Default search algo** | `best-of-n` | `best-of-n` | `single-pass` |\n| **Reward model** | AF2 (`af2folding`) | RF3 (`rf3folding`) | `null` (no reward at default) |\n| **Inverse folder** | `soluble_mpnn` | `ligand_mpnn` | `ligand_mpnn` |\n| **Default refold backend** | `colabdesign` (AF2) | `rf3_latest` | `rf3_latest` |\n| **Analysis `result_type`** | `protein_binder` | `ligand_binder` | `motif_ligand_binder` |\n| **Required ckpts (`complexa download`)** | `--complexa --all` (AF2 in community) | `--complexa-ligand --all` (RF3 in community) | `--complexa-ame --all` (RF3 in community) |\n\nFor the full per-pipeline breakdown (reward weights, success thresholds,\nanalysis modes), see [reference/pipelines.md](reference/pipelines.md).\n\n## Step 3: Gather parameters\n\nUse AskUserQuestion to fill in the four parameters that vary every run. Default\nto sensible production settings if the user has no preference.\n\n- **Target name** — must be a key in the relevant dict (`targets_dict.yaml` for\n  protein binder, `ligand_targets_dict.yaml` for ligand, `ame_dict_v2.yaml` for\n  AME). If the user names a target that is not in the dict, hand off to\n  `complexa-target` to add it first.\n- **Run name** — a short identifier appended to the output dir (e.g. `pdl1_v1`).\n- **Search algorithm** — default to `beam-search` with `beam_width=8` and\n  `n_branch=4` for production. Use `single-pass` for a quick smoke test.\n- **Evaluation refold backend** — protein binder defaults to `colabdesign`\n  (AF2); ligand/AME default to `rf3_latest`. Use `esmfold` for fast iteration\n  (worse but seconds per sample).\n\n## Step 4: Validate\n\nValidate before running. This is cheap (seconds) and catches missing ckpts,\nmissing env vars, unknown override keys, and missing target entries — all of\nwhich would otherwise abort the pipeline mid-evaluation after hours of\ngeneration.\n\n```bash\ncomplexa validate design configs/search_binder_local_pipeline.yaml \\\n    ++generation.task_name=02_PDL1 \\\n    ++metric.binder_folding_method=colabdesign\n```\n\nThe validator returns non-zero on failure and prints a pass/fail report.\nRe-run the command with the suggested overrides until it returns clean.\n\n## Step 5: Run the pipeline\n\n`complexa design` is the right tool for the full 4-stage run — it orchestrates\n`generate → filter → evaluate → analyze` as sequential subprocesses with a\nshared run name, log dir, and multi-GPU split (see `run_design_pipeline` in\n`src/proteinfoundation/cli/cli_runner.py`). Re-implementing that manually\nloses the per-stage log routing and progress prints.\n\nUse `++` (forced) Hydra overrides; they apply to all stages. The minimal\nproduction protein-binder invocation:\n\n```bash\ncomplexa design configs/search_binder_local_pipeline.yaml \\\n    ++run_name=pdl1_v1 \\\n    ++generation.task_name=02_PDL1 \\\n    ++generation.search.algorithm=beam-search \\\n    ++generation.search.beam_search.beam_width=8 \\\n    ++metric.binder_folding_method=colabdesign\n```\n\nFor ligand binder / AME, swap the pipeline YAML and target name per Step 2's\ncheat sheet — every other override above is pipeline-agnostic and can be\nreused as-is.\n\nAdd `--verbose` to stream logs to the terminal instead of `./logs/`. The skill\ndoes not poll progress — the user re-invokes if they want a status; point them\nat `complexa status` and `./logs/design_pipeline_*/`.\n\n### Direct module invocation (debug fallback)\n\nTo debug a single stage without pipeline orchestration, invoke the underlying\nHydra module directly. `complexa generate CONFIG` is just a logged subprocess\nwrapper around this:\n\n```bash\npython -m proteinfoundation.generate \\\n    --config-path \"$(realpath configs)\" \\\n    --config-name search_binder_local_pipeline \\\n    ++run_name=debug_pdl1 \\\n    ++generation.task_name=02_PDL1\n```\n\nSame pattern for `proteinfoundation.{filter,evaluate,analyze}`. Use this when\nyou want to attach `ipdb`, run under `nsys`, or skip the pipeline log dir.\nPrefer `complexa generate/filter/evaluate/analyze` for normal one-shot runs\n(you get logging and parallel job splitting for free).\n\n**AME-specific gotcha (when running AME with RF3 refold)**: RF3 will try to\ncomplete missing atoms on the ligand based on its CCD code, which produces\nshape errors in RMSD calculations and the wrong structure. Before RF3 sees the\nPDB, rename the ligand residue to `L:0` so RF3 treats it as a generic ligand\nand skips atom completion. If your AME generation outputs already encode the\nligand this way (the canonical Complexa pipeline does), no extra step is\nneeded; otherwise patch each PDB with `atomworks.io`:\n\n```python\nfrom atomworks.io import load_any, to_pdb_file\natom_array = load_any(\"my_design.pdb\")[0]\nligand_mask = atom_array.chain_id == \"A\"\natom_array.res_name[ligand_mask] = \"L:0\"\nto_pdb_file(atom_array, \"my_design_rf3_ready.pdb\")\n```\n\nSame applies if you're piping AME outputs into the `complexa-evaluate-pdbs`\nskill with an RF3 backend. Skip the rename for AF2/colabdesign refold or for\nnon-AME pipelines.\n\nWall-clock at default (`nsteps=400`, `beam_width=8`, `batch_size=16`, 100\ndesigns, colabdesign eval) is ~30–120 minutes on a single A100/H100.\n\n## Step 6: Collect results\n\nOutputs land in two directories. Surface both:\n\n```bash\nls ./inference/${CONFIG_STEM}_${TASK}_*${RUN_NAME}/   # generated PDBs + filter\nls ./evaluation_results/${RUN_NAME}/                  # per-design CSV + analysis\n```\n\nRead the combined results CSV and summarize:\n\n```bash\nls ./evaluation_results/*/binder_results_*_combined.csv\nls ./evaluation_results/*/motif_binder_results_*_combined.csv  # AME\nls ./evaluation_results/*/res_filter_*_pass_*.csv              # success rate\nls ./evaluation_results/*/res_div_foldseek_*.csv               # FoldSeek diversity\n```\n\nPull the success rate from `res_filter_binder_pass_*.csv`, the per-design\nmetrics (interface pAE, pLDDT, scRMSD) from the combined CSV, and FoldSeek\nTM-score diversity from `res_div_foldseek_*.csv`. Report top-N designs by\ni_pAE (protein binder) or min_ipAE (ligand binder).\n\n## Step 7: Emit manifest\n\nDrop a JSON manifest beside the results so the run is replayable. The shared\nhelper captures the command, config, git SHA, and pointers to the result CSVs.\n\n```bash\npython3 .claude/skills/_shared/scripts/write_manifest.py \\\n    --output-dir ./evaluation_results/${RUN_NAME} \\\n    --command \"complexa design configs/search_binder_local_pipeline.yaml ++run_name=${RUN_NAME} ++generation.task_name=${TASK}\" \\\n    --skill complexa-design \\\n    --out ./run_manifest.json\n```\n\nSurface the manifest path and the result CSV to the user.\n\n## Most-common overrides\n\nThe 10 overrides that cover ~90% of runs. Full reference (every key, type,\ndefault) is in [reference/overrides.md](reference/overrides.md).\n\n| Override | Default | What it controls |\n|----------|---------|------------------|\n| `++generation.task_name=<name>` | (per config) | Which target / AME task to design for |\n| `++run_name=<str>` | (config stem) | Output dir suffix and CSV tag |\n| `++generation.search.algorithm=beam-search` | `best-of-n` (binder/ligand), `single-pass` (AME) | Search strategy |\n| `++generation.search.beam_search.beam_width=8` | `4` | Beam-search width (more = better designs, slower) |\n| `++generation.args.nsteps=200` | `400` | Diffusion steps (fewer = faster, lower quality) |\n| `++generation.dataloader.batch_size=8` | `16` (binder/ligand/AME) | Drop to 8 on a 40GB GPU |\n| `++generation.filter.filter_samples_limit=500` | `1000` | Top-N samples to keep after filtering |\n| `++metric.binder_folding_method=esmfold` | `colabdesign` (binder), `rf3_latest` (ligand/AME) | Evaluation refold backend |\n| `++metric.num_redesign_seqs=8` | `2` | ProteinMPNN/LigandMPNN/SolubleMPNN sequences per design |\n| `++aggregation.success_thresholds.i_pAE.threshold=10.0` | `7.0` (protein binder) | Loosen / tighten success criteria |\n\n## Hardware requirements\n\n| Resource | Minimum | Recommended |\n|----------|---------|-------------|\n| GPU | 1x CUDA GPU, 40 GB VRAM | A100 / H100 / L40S, 80 GB VRAM |\n| CPUs | 16 | 24 (the `ncpus_` default in every pipeline config) |\n| Disk | 50 GB at `./inference/` + `./evaluation_results/` | 200 GB for sweep runs |\n| RAM | 32 GB | 64 GB+ |\n\nTypical wall-clock for 100 designs, `beam_width=8`, default `nsteps=400`:\n\n- Protein binder + colabdesign refold: ~60–120 min on 1x A100/H100.\n- Ligand binder + RF3 refold: ~90–180 min (RF3 dominates).\n- AME + RF3 refold: ~120–240 min.\n- Any pipeline + ESMFold refold: ~30–60 min (fast iteration).\n\nBumping `gen_njobs=2` and `eval_njobs=2` halves wall-clock on a 2-GPU host. See\n`.claude/skills/_shared/reference/hardware.md` for per-pipeline VRAM tables.\n\n## Troubleshooting (common cases)\n\n| Symptom | Cause | Fix |\n|---------|-------|-----|\n| `CUDA out of memory` in generate | `batch_size: 16` too big on 40GB GPU | `++generation.dataloader.batch_size=8` |\n| `CUDA out of memory` in evaluate | AF2 / RF3 batched too aggressively | `++eval_njobs=1` and `++metric.num_redesign_seqs=2` |\n| `InterpolationKeyError: AF2_DIR` | colabdesign eval but `.env` does not set `AF2_DIR` | Set `AF2_DIR` in `.env` or `++metric.binder_folding_method=esmfold` |\n| `InterpolationKeyError: RF3_CKPT_PATH` | RF3 eval but RF3 not installed | `complexa download --all` or switch eval backend |\n| `KeyError: 'task_name' not in target_dict_cfg` | Target not in `targets_dict.yaml` / `ligand_targets_dict.yaml` / `ame_dict_v2.yaml` | Use `complexa-target` skill to add it |\n| 0 designs pass success thresholds | Defaults too strict for this target | Loosen via `++aggregation.success_thresholds.*` |\n\nFor the full list (chain-ID mismatches, hotspot residues, ligand residue\nrenaming for RF3, missing inverse-folding models, etc.) see\n[reference/troubleshooting.md](reference/troubleshooting.md).\n\n---\n\nFor per-pipeline details (which model, reward, inverse folding, evaluation\nbackend, `result_type`, LoRA settings, `USE_V2_COMPLEXA_ARCH` toggle), see\n[reference/pipelines.md](reference/pipelines.md).\n\nFor the full override reference (every `generation.*`, `metric.*`,\n`aggregation.*` key with type, default, example, and effect), see\n[reference/overrides.md](reference/overrides.md).\n\nFor all troubleshooting cases (cause + fix + source), see\n[reference/troubleshooting.md](reference/troubleshooting.md).\n"
}

SHA-256: c0371904d8913e7cd3bd1933a7d2f48fc1b77bd5fc770f87724f02e7a5e3a361