← Files Biohub ESMARCHIVED FILE

skills/esm-atlas/references/api.md

12.1 KB · Oct 5, 2026 · 18:31 UTC

↓ Download file

# ESM Atlas v1alpha1 HTTP contract

Primary official sources:

- [API reference](https://biohub.ai/esm/protein/atlas/api-docs/api_reference.html)
- [OpenAPI 3.1 schema](https://biohub.ai/esm/protein/atlas/api-docs/_static/openapi.json)
- [similarity-search example](https://biohub.ai/esm/protein/atlas/api-docs/examples/sae_search.html)
- [protein and cluster example](https://biohub.ai/esm/protein/atlas/api-docs/examples/protein_lookup.html)
- [feature examples](https://biohub.ai/esm/protein/atlas/api-docs/examples/feature_browse.html)

Verified against source code and all linked reference, example, concept, FAQ, and Swagger sections on 2026-07-31. Base URL: `https://biohub.ai`. Atlas currently requires no client-side authentication. It is an alpha API, so preserve the complete decoded response and validate its schema before deriving artifacts.

## Published operations

| Method and path | Request | Success contract |
| --- | --- | --- |
| `GET /esm/protein/api/v1alpha1/features` | no parameters | `200` JSON containing the complete 16,384-item feature catalog |
| `GET /esm/protein/api/v1alpha1/features/{feature_index}` | path index 0–16,383 | `200` JSON feature detail; `422` validation error |
| `GET /esm/protein/api/v1alpha1/clusters/{protein_hash}` | representative MD5; optional `topk_features` 1–100, default 10 | `200` JSON cluster detail; `422` validation error |
| `GET /esm/protein/api/v1alpha1/proteins/{protein_hash}` | MD5 plus protein query fields below | `200` JSON protein record; `422` validation error |
| `GET /esm/protein/api/v1alpha1/proteins/{protein_hash}/thumbnail/{thumbnail_type}` | type `plddt` or `pct-characterized` | `200` PNG, immutable for one year; `422` validation error |
| `GET /esm/protein/api/v1alpha1/similarity-search` | sequence plus search fields below | `200` JSON search response; `422` validation error |
| `POST /esm/protein/api/v1alpha1/proteins/batch` | JSON batch request below | small batch: `200` ZIP; large batch: `202` JSON job handle |
| `GET /esm/protein/api/v1alpha1/proteins/batch/jobs/{job_id}` | job ID | `202` pending, `200` terminal/download-ready, `410` expired |
| `DELETE /esm/protein/api/v1alpha1/proteins/batch/jobs/{job_id}` | job ID | idempotent `204`; unknown job may return `404` |

## Similarity search

The required `sequence` is accompanied by optional `topk_results` 1–100 (default 10), `topk_features` 1–100 (default 20), `min_similarity` 0–1 (OpenAPI default 0.5), `cluster_pct_characterized_max` 0–100 or null, and `include_cluster_info` (default false). `cluster_pct_characterized_max=0` is only the documented proxy for clusters with no characterized Pfam annotations.

The raw endpoint accepts 1 to 2,048 sequence characters, uppercases input, and accepts `ACDEFGHIKLMNPQRSTVWYXBUZO|:`; separators count toward the limit. The published worked example still states an 800-residue maximum despite the 2,048-character schema/backend limit. The bundled plugin therefore caps searches at 800 and accepts `ACDEFGHIKLMNPQRSTVWYXBUZO`, not the raw-only `|` and `:` separators. Never clip silently; for longer inputs, ask for a domain or search labeled windows.

`SAESearchResponseV2` requires `query_sequence` and `similar_proteins`; it may also contain the query `protein_hash`, `top_features_across_results`, and `restricted_count`. Every `SimilarProtein` requires exactly these named fields:

- `protein_hash`: lowercase MD5
- `protein_accession`: source-prefixed accession when available, otherwise the protein hash
- `sequence_length`: nonnegative integer; 0 means metadata was unavailable. The bundled client currently rejects this fallback, so a hit carrying `sequence_length: 0` surfaces as a local validation failure, not a malformed Atlas response. The published field is not `protein_length`.
- `similarity_score`: cosine similarity from 0 to 1

Optional hit fields are `pdb`, `ptm`, `mean_plddt`, `residues_plddt`, `cluster_size`, and `protein_name`. Cluster size and name are populated only when `include_cluster_info=true`. `top_features_across_results` entries require `feature_index`, `occurrence_count`, `min_activation`, `max_activation`, and `mean_activation`. `restricted_count` is the count withheld by Atlas filtering.

`similarity_score` is cosine similarity over normalized, max-pooled SAE vectors, not sequence identity. Preserve Atlas order, report relevant score spread, and do not call a narrow range tied or noisy without a supported tolerance. Exact self-matches need not score 1.

## Protein lookup

Query fields are `topk_features` 1–100 (default 10), `fold_on_miss` (provider default true), `normalize_features` (default true), and repeated integer `feature_indices` capped at 100. The plugin always sends `fold_on_miss=false` unless the caller opts in, and rejects an explicit empty `feature_indices` list because an omitted query means top-K mode. Nonempty `feature_indices` returns those features in request order, uses value 0 for no recorded activation, and skips indices outside the catalog; clients must not reject those documented out-of-catalog values before Atlas sees them. The client then requires the returned `sae_features` indices to equal the requested in-catalog indices in caller order. Without explicit indices, it rejects a response containing more than `topk_features` entries.

The endpoint is keyed by a hash for a protein record already known to Atlas. An opt-in fold is a guarded two-step operation: first request `fold_on_miss=false`; return that record immediately when it already contains coordinates; otherwise require the actual provider-returned sequence, verify that it matches the requested MD5, validate the supported Atlas alphabet, and enforce the plugin's 699-sequence-character limit before sending `fold_on_miss=true`. `sequence_length` alone, missing sequence evidence, a hash mismatch, an unsupported character, or an over-limit sequence stops before the folding request. Atlas folds its stored protein and does not require or accept a caller-supplied sequence through this route. The published Atlas overview describes folding as less than 700 residues, while the backend's unlisted caller-sequence `/fold` route accepts 1 to 700 characters; the plugin retains the documented 699 maximum and does not expose that route, so do not call it ad hoc from this skill. Use similarity search first when starting from a sequence. The response requires only `protein_hash`; documented optional fields are `header`, `source`, `accession`, `sequence`, `sequence_length`, `ptm`, `mean_plddt`, `residues_plddt`, `sae_features`, `protein_activations`, `per_residue_activations`, `pdb`, `cluster_rep_protein_hash`, and `folded_on_demand`.

`sae_features` entries require `feature_index`, `label`, `description`, and `value`; each residue region requires `start`, `end`, `peak_residue`, and `mean_activation`. Sparse activations use COO objects with `indices`, `values`, and `shape`; every coordinate must be within its axis extent, and complete coordinate tuples must be unique. Similarity-search hits are representatives and the raw Atlas cluster endpoint accepts their `protein_hash`. The bundled guarded multi-step workflows still require a non-null provider-returned `cluster_rep_protein_hash` and do not infer it from a hit; for arbitrary protein lookups, use that returned representative hash.

For `esmc-6b-2024-12-sae-layer60-k64-codebook16384`, k64 is the SAE sparsifier, not API truncation. It leaves at most 64 positive activations per residue, and the returned COO entries are the complete stored post-sparsification representation. Omitted coordinates are zero in that representation; uint8 storage quantization can round small positive values to zero. The protein-level `sae_features` selection is separately governed by `topk_features`.

## Cluster lookup

`ClusterDetailResponse` requires `protein_hash`, `cluster_size`, `cluster_pct_characterized`, `cluster_mean_domain_coverage`, and `member_protein_hashes`. Optional metadata includes `protein_name`, `source`, `accession`, Pfam-domain mappings, representative SAE features, LCA taxonomy (`rank`, `name`), `top_phyla`, `mean_plddt`, and `ptm`. Each Pfam value requires `count` and `name`.

## Features

The list endpoint has no pagination or filtering and returns all 16,384 entries in `data`. Every item requires `feature_index`, `label`, and `description`.

Feature detail requires `feature_index`, `label`, `summary`, `description`, `uniref90_frequency`, `uniref90_idf`, `uniref90_max_activation`, and `threshold`. Optional/defaulted fields are `activation_pattern`, `category`, `exemplar_protein_families`, `top_100_uniref_ids`, `top_swissprot_activations`, and `decoder_nearest_neighbors`. UniRef/SwissProt list entries pair their identifier with a numeric `activation`. The list endpoint's `description` currently contains the short summary for backward compatibility; detail `summary` and longform `description` are distinct.

## Thumbnails

The endpoint returns `image/png` for `plddt` and `pct-characterized`, with `Cache-Control: public, max-age=31536000, immutable` and `ETag: "{protein_hash}-{thumbnail_type}"`. The ETag identifies the requested resource and is not a checksum of the PNG bytes. The generated OpenAPI file incorrectly labels the empty `200` schema as `application/json`; validate the returned PNG bytes.

## Batch lifecycle

`BatchProteinRequest` requires a nonempty `protein_hashes` list and deduplicates it to at most 500 entries. The raw endpoint reports malformed hashes as per-protein `invalid_hash` errors; the plugin validates every entry locally and rejects the entire input when any hash is malformed. Defaults are `topk_features=10`, `include_structure=true`, `include_cluster_info=true`, `include_sequence=true`, and both protein-level and per-residue feature data. The property description explicitly allows `include_features` to be `false` (omit features), `true` (include both), or an object with `protein_level` and/or `per_residue`. Neither object sub-flag is required; an omitted sub-flag defaults to true, so `{}` is equivalent to enabling both. Preserve boolean and partial-object wire forms rather than inventing required keys.

The endpoint narrative defines a `200` ZIP for a small synchronous batch and a `202` `BatchProteinResponse` for an asynchronous batch, even though the generated OpenAPI response table incompletely lists only `200 application/json`. The job response requires only `status`, whose value is one of `pending`, `completed`, `cancelled`, `failed`, or `expired`; `job_id`, `poll_url`, `download_url`, `created_at`, `completed_count`, and `total_count` are nullable/optional. A `202` submission still needs a usable `job_id` to be actionable, but a later status response may omit it because the requested path already identifies the job. Pending polling uses HTTP 202. A completed response uses HTTP 200 and supplies a time-limited `download_url`; an expired job uses HTTP 410.

Cancellation is idempotent and returns 204 for pending, already completed, or already cancelled jobs, and 404 for an unknown job. Completed output remains available. Preserve job state, reacquire ephemeral download URLs by polling, and validate the completed ZIP before publication.

## Confidence scaling and schema conflicts

Atlas JSON exposes `ptm`, `mean_plddt`, and `residues_plddt` on a 0–1 scale. Embedded PDB B-factors carry residue-level pLDDT on the same 0–1 scale; pTM and mean pLDDT are not B-factor fields. The public concepts page explains pLDDT on the familiar 0–100 scale; multiply by 100 only for conventional presentation and preserve the original wire scale.

The generated `BatchProteinRequest` schema resolves `include_features` only to the object schema even though that property's own description explicitly supports booleans. The other known documentation inconsistencies are the 800-vs-2,048 search length and incorrect OpenAPI media/status declarations for thumbnails and batch delivery. Handle those explicitly as above; do not invent aliases for documented JSON fields.

## Index coverage

Similarity search indexes cluster representatives only. Other Atlas proteins remain retrievable by hash but do not appear as search hits. Missing known relatives can therefore reflect index coverage rather than an absence of homologs. Do not infer search eligibility from a null `cluster_rep_protein_hash`.

SHA-256: 39454f9e6eecd94656a88d64c5018dd9c00151e6b65b81bf6d727f4085a05b67