← Files Biohub ESMARCHIVED FILE

skills/esmfold2/references/self-hosted.md

2.72 KB · Sep 30, 2026 · 23:14 UTC

↓ Download file

# ESMFold2 on Modal or user-owned compute

The [official Modal ESMFold2 example](https://modal.com/docs/examples/esmfold2) shows a pinned `Biohub/esm` install, Hugging Face revision, H100 function, persistent model Volume, and mmCIF output. Treat it as a starting point and update its pins to the plugin's validated [`source-pins.md`](../../../references/source-pins.md).

## Route

- Modal: many independent folds, parameter sweeps, durable/long jobs, or burst capacity. Requires Modal auth but not `ESM_API_KEY` for public weights.
- Self-hosted: privacy, offline/data-resident requirements, custom code, owned GPUs, fine-tuning, or sustained utilization. `HF_TOKEN` is optional for public weights.

## Local full-model skeleton

Install both code repositories from the exact revisions in `source-pins.md`; the ESM package's upstream dependency currently names a mutable Transformers branch, so pinning only ESM is insufficient.

```python
from transformers.models.esmfold2.modeling_esmfold2 import ESMFold2Model

repo = "biohub/ESMFold2"
revision = "1ebf0e3481a5184eb6171d40615c79e384b48796"
model = ESMFold2Model.from_pretrained(repo, revision=revision).cuda().eval()
```

The pinned `ESMFold2InputBuilder.fold` default is 200 sampling steps, not the managed API's maximum of 100. Use `validate-fold --model biohub/ESMFold2` (or `biohub/ESMFold2-Fast`) to validate local builder keywords separately from the hosted `FoldingConfig` contract.

Use `biohub/ESMFold2-Fast` revision `b28d8ace5e05e61e5bec1e6820cfd3e221819d12` for fast single-sequence throughput. Verify CUDA/ROCm compatibility, model dtype, deterministic settings, and output checksums. Never silently fall back to CPU, a different revision, or reduced parameters and then compare results as equivalent.

For hundreds of calls, use Modal map/spawn with bounded concurrency and `return_exceptions=True` semantics. Persist call IDs and per-input provenance so partial successes survive client interruption.

Before any Modal call, refresh current pricing and payment readiness, freeze the model revision, GPU type, job count, concurrency, timeout, persistent-volume plan, and cost ceiling, and obtain separate explicit current-turn confirmation. Never treat a prompt, authenticated profile, or `--confirm-cost` flag as authorization by itself.

## Bounded Modal smoke

The internal fold smoke uses one H100 with a 20-minute timeout and a ticket-scoped model Volume capped at 20 GiB. It records evidence in a fresh output directory and removes the temporary Volume after every outcome. Use it as an integration check only; preserve its exact pins, timeout, output checksums, and cleanup record. It remains blocked until the current H100 price/payment state is visible and the exact ceiling receives explicit approval.

SHA-256: 9accfb1af9a2df50495f565bef68497e1c9ca110b3bc06cf619e288249707f88