{"id":17599,"plugin_id":"plugins_6a76572d8f8081918362aa7ff90947fb","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.757Z","digest":"965092786ad88f5aeea04643c73670c31899c832e5cdd17c74732375484de6c7","against":null,"payload":{"name":"msa-structure-prediction-pipeline","description":"Run a complete protein structure prediction pipeline using NVIDIA BioNeMo NIMs: search for MSA alignments with MSA-Search (ColabFold), then predict the structure with OpenFold3 using the retrieved alignments. Use this skill whenever the user wants to predict a protein structure with maximum accuracy using MSA context, run the full AlphaFold3-style pipeline, generate MSA-informed structure predictions, or improve structure prediction accuracy by providing evolutionary information. Triggers on: MSA structure prediction pipeline, structure prediction pipeline, MSA-informed prediction, OpenFold3, ColabFold MSA, AlphaFold3 pipeline, protein structure, homology search, a3m alignment, UniRef30, NIM microservice. This pipeline chains MSA-Search and OpenFold3.","included_files":[],"skill_md_contents":"---\nname: msa-structure-prediction-pipeline\ndescription: >\n  Run a complete protein structure prediction pipeline using NVIDIA BioNeMo NIMs:\n  search for MSA alignments with MSA-Search (ColabFold), then predict the structure\n  with OpenFold3 using the retrieved alignments. Use this skill whenever the user wants\n  to predict a protein structure with maximum accuracy using MSA context, run the\n  full AlphaFold3-style pipeline, generate MSA-informed structure predictions, or\n  improve structure prediction accuracy by providing evolutionary information.\n  Triggers on: MSA structure prediction pipeline, structure prediction pipeline, MSA-informed prediction, OpenFold3,\n  ColabFold MSA, AlphaFold3 pipeline, protein structure, homology search, a3m alignment,\n  UniRef30, NIM microservice. This pipeline chains MSA-Search and OpenFold3.\nlicense: Apache-2.0 AND CC-BY-4.0\nallowed-tools: Bash, Read, Write, AskUserQuestion\n---\n\n# MSA Structure Prediction Pipeline\n\nPredict protein structures with high accuracy by chaining two BioNeMo NIMs:\n\n```\nStep 1: MSA-Search  →  Step 2: OpenFold3\n(Search homologs)       (Predict structure with MSA)\n```\n\n---\n\n## Overview\n\nWhy chain these NIMs?\n- **MSA-Search** finds evolutionary homologs in UniRef30 and ColabFold databases using GPU-accelerated MMSeqs2. The resulting alignment provides crucial evolutionary information.\n- **OpenFold3** uses the MSA to improve structure prediction accuracy — especially for sequences where no close homolog exists in PDB.\n- Running MSA-Search first means OpenFold3 gets the full evolutionary context rather than a single-sequence prediction.\n\n---\n\n## Before you start\n\nConfirm with the user:\n1. **Query sequence**: amino acid sequence to predict\n2. **MSA depth**: how many sequences to retrieve (default 500; more = slower but more context)\n3. **API mode**: hosted or local Docker?\n\nNote: local MSA-Search requires 1.4 TB of database storage — strongly recommend hosted unless the user has that infrastructure.\n\nFor local Docker, do not assume MSA-Search and OpenFold3 are both on\n`localhost:8000` concurrently. Run one container at a time and hand off the A3M\nfile, or start each NIM on a distinct host port and set the URLs explicitly.\n\n---\n\n## Step 1: Search for MSA with MSA-Search\n\n```python\nimport requests, json, os\nfrom pathlib import Path\n\nNGC_API_KEY = os.environ[\"NGC_API_KEY\"]\nHOSTED = True\n\nquery_sequence = \"<YOUR_PROTEIN_SEQUENCE>\"\n\nif HOSTED:\n    msa_url = \"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict\"\n    headers = {\"Content-Type\": \"application/json\",\n               \"Authorization\": f\"Bearer {NGC_API_KEY}\"}\nelse:\n    msa_url = \"http://localhost:8000/biology/colabfold/msa-search/predict\"\n    headers = {\"Content-Type\": \"application/json\"}\n\npayload = {\n    \"sequence\": query_sequence,\n    \"databases\": [\"Uniref30_2302\", \"colabfold_envdb_202108\"],\n    \"e_value\": 0.0001,\n    \"output_alignment_formats\": [\"a3m\"],\n}\n\nr = requests.post(msa_url, headers=headers, json=payload)\nr.raise_for_status()\nmsa_result = r.json()\n\n# Extract the A3M alignment\na3m_alignment = msa_result[\"alignments\"][\"Uniref30_2302\"][\"a3m\"][\"alignment\"]\n\n# Save for reference\nwith open(\"query_msa.a3m\", \"w\") as f:\n    f.write(a3m_alignment)\n\n# Count sequences in alignment\nn_seqs = a3m_alignment.count(\">\")\nprint(f\"Step 1 complete: found {n_seqs} homologous sequences\")\nprint(f\"MSA saved to query_msa.a3m\")\n```\n\n---\n\n## Step 2: Predict structure with OpenFold3\n\nPass the MSA directly into OpenFold3's `msa` field:\n\n```python\nif HOSTED:\n    of3_url = \"https://health.api.nvidia.com/v1/biology/openfold/openfold3/predict\"\nelse:\n    of3_url = \"http://localhost:8000/biology/openfold/openfold3/predict\"\n\n# Build the OpenFold3 MSA structure from the retrieved alignment\nmsa_data = {\n    \"uniref30\": {\n        \"a3m\": {\n            \"alignment\": a3m_alignment,\n            \"format\": \"a3m\"\n        }\n    }\n}\n\n# Optionally also include colabfold_envdb alignment if requested\n# env_alignment = msa_result[\"alignments\"][\"colabfold_envdb\"][\"a3m\"][\"alignment\"]\n# msa_data[\"colabfold_env\"] = {\"a3m\": {\"alignment\": env_alignment, \"format\": \"a3m\"}}\n\npayload = {\n    \"inputs\": [{\n        \"input_id\": \"prediction_with_msa\",\n        \"output_format\": \"pdb\",\n        \"molecules\": [\n            {\n                \"type\": \"protein\",\n                \"sequence\": query_sequence,\n                \"diffusion_samples\": 1,\n                \"msa\": msa_data\n            }\n        ]\n    }]\n}\n\nr = requests.post(of3_url, headers=headers, json=payload, timeout=300)\nr.raise_for_status()\nresult = r.json()\n\noutput = result[\"outputs\"][0]\nfor i, sample in enumerate(output[\"structures_with_scores\"]):\n    fmt = sample[\"format\"]\n    filename = f\"predicted_structure_{i+1}.{fmt}\"\n    with open(filename, \"w\") as f:\n        f.write(sample[\"structure\"])\n    print(f\"\\nStep 2 complete: {filename} saved\")\n    print(f\"  Confidence:  {sample['confidence_score']:.4f}\")\n    print(f\"  pLDDT:       {sample['complex_plddt_score']:.4f}\")\n    print(f\"  pTM:         {sample['ptm_score']:.4f}\")\n```\n\n---\n\n## Comparing single-sequence vs MSA-informed prediction\n\nIf the user wants to see the impact of MSA, run OpenFold3 twice — once with the full MSA and once with just the query sequence as a minimal alignment:\n\n```python\n# Minimal MSA (single sequence — same as no MSA context):\nminimal_msa = {\n    \"main\": {\n        \"a3m\": {\n            \"alignment\": f\">query\\n{query_sequence}\",\n            \"format\": \"a3m\"\n        }\n    }\n}\n```\n\nA larger, higher-quality MSA typically yields higher pLDDT and lower pDE, especially for proteins with many known homologs.\n\n---\n\n## For protein complexes\n\nUse the `/paired/predict` endpoint of MSA-Search to get paired alignments for multi-chain complexes, then pass each chain's alignment into the corresponding molecule's `msa` field and `paired_msa` fields:\n\n```python\n# Paired MSA search endpoint for complexes:\nmsa_paired_url = \"https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict\"\npaired_payload = {\n    \"sequences\": [chain_A_sequence, chain_B_sequence],\n    \"e_value\": 0.0001,\n}\n```\n\n---\n\n## Quick reference — skill dependencies\n\n| Step | Skill | Key endpoint |\n|---|---|---|\n| MSA search | `msa-search-nim` | `/biology/colabfold/msa-search/predict` |\n| Structure prediction | `openfold3-nim` | `/biology/openfold/openfold3/predict` |\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}