← NVIDIA BioNeMo Agent ToolkitCONTENT HISTORY

Update to NVIDIA BioNeMo Agent Toolkit

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "openfold2-nim",
  "description": "Use this skill for OpenFold2, NVIDIA's BioNeMo NIM microservice for monomer protein structure prediction. Invoke whenever the user mentions OpenFold2, AlphaFold2-like monomer folding, protein sequence-to-structure prediction, A3M MSAs, mmCIF templates, hosted NVIDIA API calls, or local Docker deployment.",
  "included_files": [
    {
      "relative_path": "references/api.md",
      "size_in_bytes": 4543
    },
    {
      "relative_path": "references/examples.md",
      "size_in_bytes": 1884
    },
    {
      "relative_path": "references/parameters.md",
      "size_in_bytes": 1852
    },
    {
      "relative_path": "references/science.md",
      "size_in_bytes": 2949
    },
    {
      "relative_path": "references/validation.md",
      "size_in_bytes": 1516
    }
  ],
  "skill_md_contents": "---\nname: openfold2-nim\ndescription: >\n  Use this skill for OpenFold2, NVIDIA's BioNeMo NIM microservice for monomer protein structure prediction. Invoke whenever the user mentions OpenFold2, AlphaFold2-like monomer folding, protein sequence-to-structure prediction, A3M MSAs, mmCIF templates, hosted NVIDIA API calls, or local Docker deployment.\nlicense: Apache-2.0 AND CC-BY-4.0\ncompatibility: \"requests>=2.28\"\nallowed-tools: Bash, Read, Write, AskUserQuestion\n---\n\n# OpenFold2 NIM\n\nPredict a single protein-chain structure from an amino-acid sequence, with\noptional A3M multiple sequence alignments and mmCIF templates. Use this\n`SKILL.md` for basic hosted/local NIM use; load supplemental files only when\nthe task needs deeper context:\n\n- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.\n- `references/science.md`: model scope, strengths, limitations, and handoffs.\n- `references/parameters.md`: MSA, template, model-selection, and relax effects.\n- `references/validation.md`: artifact and scientific sanity checks.\n- `references/examples.md`: compact hosted/local payload patterns.\n\n## Choose Mode\n\nAsk only when context is unclear:\n\n> Hosted NVIDIA API or local Docker NIM?\n\n- Hosted URL: `https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template`\n- Local URL: `http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template`\n- Local readiness: `http://localhost:8000/v1/health/ready`\n\nMode difference: hosted and local use the same prediction path except local\ndoes not include `/v1/`. Hosted requests use `Authorization: Bearer\n$NGC_API_KEY`; local inference requests use no auth header after readiness.\n\n## Auth And Environment\n\nDo not print API keys. Confirm they exist with shell tests, not echoes.\n\nHosted needs `NGC_API_KEY` in the request header. Supported local Docker\nstartup uses `NGC_API_KEY`, or `NVIDIA_API_KEY` as a fallback, plus\n`LOCAL_NIM_CACHE`. A repo-root `.env` file may be sourced as a local override.\n\n## Local Docker\n\nUse the official OpenFold2 NIM image and mount `LOCAL_NIM_CACHE` at\n`/opt/nim/.cache`. Current docs recommend at least 80 GB disk, 64 GB system\nRAM, 8 CPU cores, and one supported GPU; the container is roughly 55 GB and\nfirst startup downloads about 10 GB of model parameters.\n\nWhen writing local setup commands, copy the preflight below exactly. Do not\ndrop `.env`, `NVIDIA_API_KEY`, `LOCAL_NIM_CACHE`, or the no-auth local request.\n\n```bash\nset -a\n[ -f .env ] && . ./.env\nset +a\n\nif [ -z \"${NGC_API_KEY:-}\" ] && [ -n \"${NVIDIA_API_KEY:-}\" ]; then\n  export NGC_API_KEY=\"$NVIDIA_API_KEY\"\nfi\n: \"${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}\"\n: \"${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}\"\n\necho \"$NGC_API_KEY\" | docker login nvcr.io --username '$oauthtoken' --password-stdin\n\nexport NIM_TEST_GPU=\"${NIM_TEST_GPU:-0}\"\nmkdir -p \"${LOCAL_NIM_CACHE}\"\nchmod 777 \"${LOCAL_NIM_CACHE}\"\n\ndocker run --rm --name openfold2 \\\n  --runtime=nvidia \\\n  --gpus \"device=${NIM_TEST_GPU}\" \\\n  -e NGC_API_KEY \\\n  -v \"${LOCAL_NIM_CACHE}:/opt/nim/.cache\" \\\n  -p 8000:8000 \\\n  nvcr.io/nim/openfold/openfold2:latest\n```\n\nReadiness check:\n\n```bash\nuntil curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done\n```\n\n## Request Pattern\n\nUse Python `requests`; curl escaping is fragile for A3M/mmCIF text. The\n`sequence` field is required. `input_id`, `alignments`, `selected_models`,\n`relax_prediction`, `use_templates`, and `explicit_templates` are optional.\n\n```python\nimport os\nimport requests\n\nhosted = True\nurl = (\n    \"https://health.api.nvidia.com/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template\"\n    if hosted\n    else \"http://localhost:8000/biology/openfold/openfold2/predict-structure-from-msa-and-template\"\n)\nheaders = {\"Content-Type\": \"application/json\"}\nif hosted:\n    headers[\"Authorization\"] = f\"Bearer {os.environ['NGC_API_KEY']}\"\n\nseq = \"MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPT\"\npayload = {\n    \"sequence\": seq,\n    \"input_id\": \"kras_fragment\",\n    \"selected_models\": [1],\n    \"relax_prediction\": False,\n    \"alignments\": {\n        \"uniref90\": {\n            \"a3m\": {\n                \"alignment\": f\">query\\n{seq}\",\n                \"format\": \"a3m\",\n            }\n        }\n    },\n}\n\nresponse = requests.post(url, headers=headers, json=payload, timeout=300)\nresponse.raise_for_status()\nresult = response.json()\n```\n\nPayload gotchas:\n\n- OpenFold2 is monomer-only. For protein-ligand, protein-DNA/RNA, or\n  multi-chain complexes, use OpenFold3 or Boltz2 instead.\n- `sequence` must use valid amino-acid IUPAC symbols.\n- Hosted API docs list sequence length 1-1000; local docs say current NIM\n  supports sequences up to 2048 residues on supported hardware.\n- A3M alignments go under `alignments` by database name, then `a3m` with\n  `alignment` and `format`. When the user needs to create or deepen an MSA,\n  hand off to `msa-search-nim` / MSA Search and map its A3M output into this\n  `alignments` shape.\n- Starting with OpenFold2 2.0.0, use `explicit_templates` with mmCIF content;\n  do not write new HHR-template examples.\n- `selected_models` chooses AlphaFold2/OpenFold parameter sets 1-5. Select one\n  or two models for smoke tests; use all five for stronger production runs.\n\n## Save And Interpret Output\n\nThe response includes one prediction per selected model, ordered by confidence.\nSave every returned structure-like text field and the full JSON response so\nfield-shape differences are auditable. Production answers should explicitly\nwrite `.pdb` or `.cif` artifacts, preserve the response JSON, and print any\nconfidence/ranking fields the service returns.\n\n```python\nfrom pathlib import Path\nimport json\n\nPath(\"openfold2_response.json\").write_text(json.dumps(result, indent=2))\n\ndef save_strings(obj, prefix=\"openfold2\"):\n    i = 0\n    if isinstance(obj, dict):\n        for key, value in obj.items():\n            if isinstance(value, str) and (\"ATOM\" in value or value.lstrip().startswith(\"data_\")):\n                i += 1\n                ext = \"cif\" if value.lstrip().startswith(\"data_\") else \"pdb\"\n                Path(f\"{prefix}_{key}_{i}.{ext}\").write_text(value)\n            elif isinstance(value, (dict, list)):\n                i += save_strings(value, f\"{prefix}_{key}\")\n    elif isinstance(obj, list):\n        for idx, value in enumerate(obj, start=1):\n            if isinstance(value, (dict, list)):\n                i += save_strings(value, f\"{prefix}_{idx}\")\n    return i\n\nsaved = save_strings(result)\nprint(f\"saved {saved} structure artifact(s)\")\n```\n\nFor production monomer runs:\n\n- Use `selected_models: [1, 2, 3, 4, 5]` unless the user requests a smoke test.\n- Use `relax_prediction: True` in Python payloads when relaxation is desired;\n  JSON examples may show `true`.\n- State the sequence length caveat: hosted API docs list 1-1000 residues, while\n  local support-matrix docs list up to 2048 residues on supported hardware.\n- If the task is a complex rather than a monomer, redirect to OpenFold3 or\n  Boltz2.\n\nTreat tiny toy sequences and single-sequence MSAs as API smoke tests, not\nquality evidence. For scientific interpretation and validation, read\n`references/science.md` and `references/validation.md`.\n\n## Troubleshooting\n\n- `401`: missing, expired, or unauthorized NGC API key.\n- `422`: invalid amino-acid characters, sequence too long, malformed A3M, bad\n  `selected_models`, or malformed mmCIF template object.\n- Local `404`: remove `/v1/` from the prediction URL.\n- Weak structures: use MSA Search to generate deeper A3M alignments and add\n  biologically relevant mmCIF templates when appropriate.\n- Local startup stalls: first run downloads parameters into `LOCAL_NIM_CACHE`.\n"
}

SHA-256: f3f721e1f3acdbae78c87d200ef309525194b2b35880a6db6ca6511ea74843b9