← Files Biological Sequence & Alignment ViewerARCHIVED FILE

skills/biological-sequence-viewer/SKILL.md

38 KB · Oct 8, 2026 · 18:02 UTC

↓ Download file

See the change to this file →

---
name: biological-sequence-viewer
description: Open, inspect, analyze, safely edit copies of, and export local biological sequences and alignments in Codex's interactive workbench. Use for FASTA/FASTQ, GenBank/EMBL, MSA formats, ABIF/SCF chromatograms, SnapGene DNA, annotations, variants, or regional read evidence.
---

# Biological Sequence & Alignment Viewer

Use this skill when a user wants to inspect local biological sequences, alignments, annotated constructs or sequencing traces visually inside Codex. For detailed UI-to-agent mappings, binary compatibility or comparisons with specialist viewers, consult the [capability matrix](../../CAPABILITY_MATRIX.md); do not claim exhaustive Jalview, IGV, SnapGene, Benchling or FastQC parity.

## Opening Workflow

The plugin contributes native file-pane preview and a model-visible inline MCP App for supported sequence and alignment files. The model should not read or copy file bytes into the conversation merely to open the viewer.

For user-provided local files:

1. Resolve the user's intended records and requested view or analysis. Prefer explicitly named paths. If the request is ambiguous, search only the relevant workspace and ask a short clarification question when multiple plausible artifacts remain. Residue-by-residue comparisons across FASTAs, including substitutions, indels, or conservation, need an alignment; comparing lengths, GC, or QC does not.
2. Confirm that the file is a supported biological sequence or alignment artifact. For an alignment comparison, prefer opening the completed, verified alignment over the raw sequence collection. When preparing a new alignment from source records, preserve source files and record identities, record the actual producer and parameters, and verify the expected rows, common column width, and preservation of each source sequence after removing alignment gaps. Separate source files can first be assembled into one derived collection using authorized workspace tools. Combining records or observing equal lengths does not establish an alignment; do not force unaligned records into Alignment mode or rename raw FASTA to an alignment suffix.
3. Call `sequence.open_from_chat` with the exact absolute local path whenever available. Omit `presentation` for the default `"full"` viewer; pass `presentation: "inline"` only when the user or onboarding flow explicitly requests minimal controls. Hidden controls remain available by moving the pointer to the viewer's top edge or focusing that edge with the keyboard; **Show toolbar** restores their pinned state in both Sequence and Alignment modes. A workspace-relative path works only when the MCP host exposes active roots. The tool opens the rich app inline beneath its tool-call row and returns a `viewerSessionId` in `structuredContent`; use that ID immediately for same-turn `sequence.control_viewer` calls instead of waiting for later model context.
4. For an alignment comparison starting from raw records in a mounted collection, query `records` through `sequence.query_viewer` as needed for exact IDs, then call `sequence.align`. A started job is not a completed alignment: follow its status while available and confirm that the same session displays the generated Alignment with the expected rows and column count. After the automatic handoff, the generated Alignment context can establish completion even if the source-side job is no longer queryable. If an authorized external workflow or fallback produces the alignment instead, open that verified output with `sequence.open_from_chat`; saving it does not change the earlier raw-record card. Complete this handoff without requiring another user prompt. If alignment or viewer control fails, report the unresolved step and distinguish independently computed results from live viewer evidence.
5. A successful open confirms session creation (`sessionReady: true`), not rendering (`viewerReady: false`). Say that the embedded viewer is opening; only say it is ready after live viewer context or an acknowledged action confirms that it loaded. For an alignment comparison, also verify the displayed artifact, Alignment mode, and expected rows before calling the comparison ready. Make the embedded viewer card that currently shows the requested artifact the only opening affordance in the reply, including when the same card transitioned from raw records to Alignment: `Click or expand the Sequence open from chat tool card above to view it. After it loads, you can use Open in side pane, or ask me to move or control the viewer.` Do not repeat or code-format the file name or path, say that the path was opened, or imply that a path or file citation is clickable; compound alignment suffixes may not become clickable and citation clicks can open raw text instead.
6. After the app opens, call `sequence.get_context` for follow-up questions about the current mode, selection, focused record, alignment column, or visible metrics. Opening or interacting with the viewer does not automatically add its state to chat.

Do not call the app-only `sequence.open` tool directly from chat. It is reserved for Codex's native file-preview host because it requires an opaque `codex-resource://` URI that only the host can mint. Use `sequence.open_from_chat` for both user-provided local files and public starter files that Codex has fetched into the authorized workspace. The legacy `sequence.acquire_public_example` tool is optional only when the host provides independently authenticated workspace roots; it is never required for a marketplace starter and is not a fallback when roots are unavailable.

If the user asks to move, send, pin, or open an already-mounted viewer in the right pane, do not call `sequence.open_from_chat` again. Use `sequence.control_viewer` with the resolved `sessionId` and action `set_display_mode`; use `fullscreen` for the side pane and `inline` for chat. The in-view **Open in side pane** and **Return to chat** buttons provide the same standard MCP App display-mode behavior.

A newly issued chat or starter handle can restore the same viewer session and bounded workbench checkpoint when its card remounts, including after the plugin server restarts, only after the active workspace root, original session, unchanged source identity, and source revision are revalidated. Re-read current viewer context after recovery; do not open a second card or replay an already completed command. Legacy handles without an authenticated session and expired, moved, changed, or unauthorized sources fail closed. If a card reports that its source expired, moved, changed, or left the active workspace, treat **Reopen from chat** as a new request to call `sequence.open_from_chat` with the prior exact path when it is still valid. If the source moved and the current path is not known, ask the user for it; do not guess a replacement or retry the expired resource URI.

### Public starter examples

Every marketplace starter names an authoritative public record that **Codex fetches before opening the viewer**. Use host-authorized research and workspace tools to retrieve only the documented official HTTPS endpoints, write the requested source and honestly Codex-authored provenance inside the actual authorized workspace, and verify every pinned accession, version, release, checksum, format, size, subset, or derived alignment before opening. Then call `sequence.open_from_chat` exactly once with the exact absolute local path of the validated source. This creates one viewer and one session even when the plugin host does not expose MCP workspace roots. Never fetch an arbitrary URL, assume a workspace file already exists, substitute a bundled/synthetic fixture, infer workspace authority from task metadata or session history, or describe Codex-authored provenance as a plugin-issued or signed receipt.

An installed and authorized Life Science Research skill may help retrieve or confirm authoritative scientific data; if it is unavailable, Codex uses its ordinary authorized network and workspace tools against the same fixed official endpoints. Neither route requires the plugin to download or publish files. Network, rate-limit, version drift, malformed data, checksum, size, cancellation, disk, or workspace errors are terminal; explain the actionable error and do not open a viewer.

The official endpoint contracts are `https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi` for NCBI Nuccore; the fixed reviewed-protein endpoints `https://rest.uniprot.org/uniprotkb/P01116.fasta`, `https://rest.uniprot.org/uniprotkb/P01111.fasta`, and `https://rest.uniprot.org/uniprotkb/P01112.fasta` for the RAS alignment; the immutable `https://ftp.ebi.ac.uk/pub/databases/Rfam/15.1/Rfam.seed.gz` release archive with exact `RF00360` record selection for the supported Rfam example; and `https://www.ebi.ac.uk/ena/portal/api/filereport` followed only by the catalog-pinned public `https://ftp.sra.ebi.ac.uk/` FASTQ path for ENA. Validate endpoint, accession, filename, and response identity before trusting returned content; never invent a replacement URL.

The exact final marketplace prompts and expected-results contract are checked in at `starter-examples.json` and documented in `STARTER_EXAMPLES.md`. The [LSC-109](https://linear.app/openai/issue/LSC-109/sequence-and-alignment-viewer-qualify-every-refreshed-starter-example) clean-host qualification is bound to the exact version, marketplace bundle, and source digest recorded in the canonical `LSC_109_QUALIFICATION.md` and `lsc-109-qualification.json` evidence. Public release requires a matching receipt; do not present historical evidence, ordinary developer checks, or deterministic checks as qualification of a changed release.

- Exact ENA prompt: `Fetch ENA DRR037765 first 500 reads to active workspace; open and report live length range, GC, Q30, and subset provenance`. Codex retrieves the exact ENA file report and FASTQ archive, requires the pinned 127,526-byte archive MD5 `81735432a6f578b332aae58cdbd95231`, validates all 967 source reads, and writes only the first 500 complete canonical records to the authorized workspace. Canonical FASTQ uses four LF-terminated lines per record: `@` plus only the first whitespace-delimited source identifier (omit header descriptions), the uppercase unwrapped sequence, a bare `+`, and the unchanged unwrapped quality string; include the final LF. Verify the 480,372-byte subset SHA-256 `46bd72991d9c9c2bf64751e88e52548d852d5fa021da4815ee6f6517a51b18b9` and write an honest Codex-authored subset receipt before calling `sequence.open_from_chat` once with its exact absolute local path. Coordinated metadata or payload drift fails closed. After the viewer mounts, report the 469–471 bp range, GC, and Q30 only from live viewer context; preserve the verified subset provenance in the answer.
- Exact Alignment prompt: `Fetch UniProt P01116/P01111/P01112 alignment; map conserved motifs to KRAS, compute distances/tree, publish Newick to workspace`. Codex retrieves the three reviewed 189-aa `SV=1` UniProtKB records from their fixed official FASTA endpoints, verifies sequence identities, digests, and RAS motifs, and uses an authorized deterministic center-star workflow in fixed KRAS/NRAS/HRAS order with match `2`, mismatch `-1`, and gap `-2`. Require the exact 3×191, 786-byte aligned FASTA SHA-256 `cb32dd89ca7855f7666fbdf3f2ff926f935b1dbc9e7f57573f884dda7e59c68f`, record the actual producer and parameters in Codex-authored provenance, and call `sequence.open_from_chat` exactly once with the aligned file's exact absolute local path. Set exact row `P01116` as reference; map KRAS residues 10–17, 30–38, 60–76, and 116–119; compare the conserved GTPase core with the divergent CAAX tails; compute a distance matrix and exploratory neighbor-joining tree. If the user requests workspace Newick, Codex extracts the exact completed `sequence.run_analysis` tree result and creates `RAS-P01116-P01111-P01112-NJ.nwk` plus an accurately Codex-authored provenance sidecar using its authorized workspace tools without overwriting. The receipt records the actual `sequence-viewer-guide-tree-v1` analysis engine, neighbor joining, uncorrected p-distance, row identity, warning, source and output digests, and Codex as publisher. Plugin workspace publication is optional only when an independently authenticated source-bound workspace capability exists. Do not present a three-leaf guide tree as publication-grade phylogeny.
- Exact NCBI prompt: `Fetch NCBI NC_001416.1 to active workspace; open it, map cI to OR1–OR3, and translate cI with code 11`. Codex retrieves exact RefSeq accession.version `NC_001416.1` from official NCBI EFetch and verifies a complete 48,502-base GenBank record with cI `complement(37227..37940)`, code 11, protein `NP_040628.1`, and OR3/OR2/OR1 annotations. Before writing an honest provenance sidecar, independently reverse-complement the cI bases, recompute code-11 translation, and validate the pinned coding-sequence and 237-aa protein digests against the source qualifier. Call `sequence.open_from_chat` once with the verified GenBank file's exact absolute local path. In Sequence mode select the exact cI CDS, confirm its reverse-strand relationship to the three operators, then call `sequence.run_analysis` for `translate` with start `37227`, end `37940`, frame `-1`, and `geneticCodeId: 11`; compare the result with the annotated translation.

### Exact starter artifact conventions

For each RAS FASTA response, validate the reviewed `>sp|accession|entryName` prefix against its catalog record and the header fields `OS=Homo sapiens`, `OX=9606`, the matching `GN`, `PE=1`, and `SV=1`. Parse field values through the next field marker or end of the header. The optional common-name suffix `(Human)` and protein-description wording are not required identity fields; do not reject a valid header because they are absent. Independently require the catalog's sequence length, digest, and motifs before opening.

Use the source and provenance paths declared in `starter-examples.json`: `codex-viewer-examples/DRR037765-first-500.fastq`, `codex-viewer-examples/human-RAS-UniProt-SV1.aln-fasta`, or `codex-viewer-examples/NC_001416.1.gb`, with `.provenance.json` appended for each source receipt. The catalog records RAS ungapped uppercase-ASCII sequence SHA-256 values under `source.records` and exact NCBI EFetch query parameters under `source.requestParameters`. Read those values before acquisition; the aligned-file digest cannot substitute for validating the three input proteins.

For RAS, choose the first longest input as center (P01116 here). Globally align the original ungapped center independently to P01111 and then P01112 with linear-gap Needleman–Wunsch: initialize the first row/column with cumulative gap penalties; break equal scores in order diagonal, up (gap in candidate), left (gap in center). Merge pairwise center-gap columns into the growing alignment in input order: consume both columns when both centers contain residues or both contain gaps; otherwise consume the gap column and insert a gap into the opposite alignment, propagating insertions through every existing row. Preserve the original row order.

Serialize the aligned FASTA as exactly these headers, each followed by one uppercase **unwrapped** aligned sequence line, preserving `-` gaps. Use UTF-8, LF after every line including the last, and no blank lines:

```text
>P01116 RASK_HUMAN GTPase KRas, UniProtKB reviewed sequence version 1
>P01111 RASN_HUMAN GTPase NRas, UniProtKB reviewed sequence version 1
>P01112 RASH_HUMAN GTPase HRas, UniProtKB reviewed sequence version 1
```

For lambda, request EFetch `db=nuccore`, `id=NC_001416.1`, `rettype=gbwithparts`, and `retmode=text`, with the catalog's `tool` and `email` values. Save the complete response bytes without newline normalization or reserialization. The qualification baseline SHA-256 is `3c624302adeeb3c00649f549903ab781b9e75bab16069ae655833d536407367f`; changed authoritative bytes require review, never patching the response to fit. Select the cI CDS carrying `NP_040628.1`, rather than its separate gene feature. The 714-nt reverse-complemented cI CDS includes its stop codon. Compare code-11 translation with the annotated 237-aa protein after removing exactly one terminal `*`, if present; preserve internal stops and hash uppercase ASCII residues with no whitespace or final LF. OR3, OR2, and OR1 are annotated as `operator-r3`, `operator-r2`, and `operator-r1` regulatory features.

RAS Newick publication uses the actual completed live tree result, saved beside the alignment as `RAS-P01116-P01111-P01112-NJ.nwk`, with `.provenance.json` appended for its receipt. Write the live Newick with one terminating semicolon and one final LF; hash the written UTF-8 bytes. Use exclusive creation for both files, fail if either already exists, and verify the written bytes and hashes. Record Codex as publisher and the actual analysis engine, parameters, row identities, source/output hashes, and exploratory warning; never copy an expected tree from this guidance as a result.

The release-bound Rfam example remains available outside the default three-prompt portfolio. Codex retrieves only the official versioned Rfam `15.1` seed archive, selects exactly one embedded Stockholm record for `RF00360`, validates its nine-sequence `SEED` alignment, records accurate source provenance, and calls `sequence.open_from_chat` once with the resulting workspace file's exact absolute local path. Treat snoZ107/R87 C/D-box conservation with row-level nuance rather than repeating the former starter claim.

After acquisition and validation, preserve the verified database, requested/resolved identifier, derivation or subset rule, safe workspace-relative artifact/provenance paths, byte length, SHA-256, and actual provenance author in the answer. Keep acquisition separate from viewer readiness: the open result means the app is opening, and only live viewer context or an acknowledged action confirms that it loaded. Point to the existing Sequence open from chat card; do not open another card to check readiness or infer scientific results from the accession or these instructions.

## Read Context On Demand

Use the session ID returned by the opening tool, or a native preview’s explicit **Ask about selection** request. Otherwise call `sequence.list_viewers`; it includes only previously referenced sessions successfully read in trusted conversation scope, not every open viewer. A manually opened viewer first needs an explicit selection-action reference. If discovery is unavailable, use an existing explicit session reference or ask the user to identify the viewer through its selection action. If several viewers match, resolve the ambiguity instead of choosing the most recently opened one.

Start with `sequence.list_documents`, choose an actual `documentId`, then use `sequence.read_document` to read it or `sequence.search_document` with a literal `query` to locate relevant text. Read the matching range in context. These general document tools support questions beyond the predefined domain-query targets; do not invent a new query name or treat the available target enum as the limit of the viewer's data.

Search is case-insensitive by default; set `caseSensitive: true` when needed. Its match snippets are previews, not complete data. For more matches, pass `offset: nextOffset` and set `expectedRevision` to the preceding response's `revision` until `eof`; use returned UTF-16 match offsets with `read_document` to inspect exact content.

Use `sequence.get_context` when a compact orientation is useful, with only the needed `sections`: `overview`, `selection`, `viewport`, `display`, `annotations`, `results`, `inventory`, `status`, or `capabilities`. Its default overview, selection, and viewport are not the whole source. Inspect freshness, readiness, source/revision, requested/effective scope, unavailable sections, and truncation before interpreting the response. `last_observed` does not establish current visible state. An unavailable selection is not an empty selection. Use `expectedRevision` when continuing from a particular snapshot, and refresh after a conflict. Keep sequence coordinates, alignment columns, reference mappings, orientation, and loaded-subset coverage distinct.

For complete literal data, call `sequence.list_documents` and use the exact returned `documentId` with `sequence.read_document`. Choose `offset`/`length` for a range or `full: true` to read as far as one response permits. Read offsets are UTF-16 code units. To read the whole document, start at `offset: 0`, append returned `text`, and repeat `full: true` with `offset: nextOffset` and the same `expectedRevision` until `eof`. `continuationReason: "transport_limit"` means more data remains; it is not a shortened document or a complete result. There is no total-document transport ceiling. Preserve original-source versus current edited/live identity and declared coverage; do not describe retained FASTQ records or loaded evidence as the whole input. Document references remain stable during ordinary renderer activity and eight-hour sleep. Command and native observation sessions expire after 24 hours without renderer activity; model reads do not renew them. Use revisions to distinguish changing content and restart after a revision conflict or new renderer. Existing parser and computation limits remain in force.

For unchanged original file bytes, inspect the `original-source.json` access metadata and use `sequence.read_source` when available. Its `offsetDecimal` is a decimal byte offset and `length` is a byte count. For the whole source, start with `full: true` and `offsetDecimal: "0"`; decode each returned `bytesBase64` and repeat `full: true` with `offsetDecimal: nextOffsetDecimal` and the exact source `expectedRevision` until `eof`. Large results report `continuationReason: "transport_limit"` and require continuation. Byte length and EOF are separate from the text-document contract. Source access may be unavailable for resource-only sessions; do not substitute current edited/retained records and call them the complete original file.

A native preview's first **Ask about selection** lazily creates a session for the same mounted viewer and its document providers. Read `sequence.get_context` with `sections: ["capabilities"]`: when the capabilities section advertises `presentationControls`, use only those `sequence.control_viewer` actions with that section's `controlExpectedRevision` as `expectedRevision`. This supports changing the existing display and selection without reopening the file. A reference without that capability remains read-only; neither reference grants new analysis, source edits, or workspace writes. After an expired native poll is observed, the next **Ask about selection** registers a fresh session reference for the displayed source. Source changes and unmounts invalidate outstanding operations; refresh after a revision conflict instead of replaying a stale control.

Use `sequence.query_viewer` for exact records, features, ranges, rows, metrics, reads, jobs, and artifacts beyond the compact context snapshot. Follow its bounded paging contract and restart from fresh state if a revision change invalidates a continuation. Reading data does not run an analysis, change selection, or publish an artifact. Use the advertised presentation controls for selection changes; new analyses and artifact publication require the existing dedicated tools and their regular command-session authority.

## Viewer Control Workflow

Use `sequence.control_viewer` in the same conversation that contains the viewer. For same-turn initialization after `sequence.open_from_chat`, pass its returned `structuredContent.viewerSessionId` as `sessionId` directly; resolve an already-open or recovered session through the context workflow above. Coordinates are 1-based.

Session creation does not confirm rendering; a live context read or an acknowledged action confirms readiness. Submit requested controls after opening without waiting for another user turn, then read context before interpreting state. Do not delegate live-viewer control to a subagent, inspect the webview through browser or DevTools automation, or search the source tree for UI state. Use the app's context and viewer tools.

- `set_toolbar_visibility`: pass `{sessionId, action: "set_toolbar_visibility", visible: true}` to show and pin the toolbar, or `visible: false` to hide it; this changes the same state as **Show toolbar** and **Hide toolbar** in both Sequence and Alignment modes. A hidden toolbar temporarily reappears on top-edge pointer hover or keyboard focus.
- `set_display_mode`: move the live viewer between `fullscreen` and `inline`; this controls the host layout, not toolbar visibility or opening `presentation`.
- `set_mode`: switch between Sequence and Alignment when the artifact supports both.
- `set_workbench_panel` and `set_workbench_disclosure`: query `workbench-panels` or paginated `workbench-disclosures` first, then use exact returned group/panel/disclosure IDs. `panel: null` closes a tool panel; expanding a registered disclosure reveals its containing panel and ancestors. Respect `applied: false` when a target is unavailable or protected pending approval/manual-copy state prevents hiding it. Never approve a user confirmation through these controls.
- `dismiss_workbench_feedback`: query `workbench-feedback` for exact registered copy/session-error IDs, then pass `feedbackId` to dismiss that feedback. The query does not return manual-copy sequence content. This action cannot approve, confirm, cancel or overwrite a source write.
- `set_sequence_record`: switch to an exact record ID, source label, or unambiguous description from the structured `sequenceRecords` context. `set_sequence_record_browser` controls search, sort, zero-based page and expansion; query `sequence-ui-state` for the applied browser state and `records` for metadata pages.
- `focus_sequence_coordinate`, `select_sequence_range`, or `select_sequence_feature`: focus coordinates, select a range, or pin an annotated feature. Prefer exact feature IDs from structured `features` context, but a unique case-insensitive feature type, label, or qualifier value such as `CDS` or `kinase domain` also works. If a biological selector is ambiguous, use the returned exact candidate ID; include `record` when the file contains multiple records. For a circular origin-spanning range, set `wraparound: true`, pass the biological start after the origin and end before the origin, and preserve the returned two segments.
- `search_sequence` and `navigate_sequence_search_hit`: search the displayed sequence, including reverse-complement matches for nucleic acid records, and move through the resulting hits.
- `set_sequence_annotation_index`: control label/type search, zero-based page and expansion of the map's annotation index. Query `features` for exact IDs, qualifiers and compound intervals outside the displayed overview; `navigate_sequence_feature` moves to the next or previous feature.
- `set_sequence_view_options`: choose a molecule-compatible residue palette, wrap width, feature/translation/quality visibility, linear/circular/split layout, forward or reverse-complement orientation, synchronized overview/detail behavior, an explicit NCBI genetic-code table, or `originRangeExpanded`. With synchronization off, overview navigation stays independent of the detailed sequence; workbench selections still use the detail selection. Re-enabling synchronization links the overview to the current detail. Map projections never change source-coordinate semantics.
- `set_chromatogram_view_options`: set `firstBase` and `basesPerWindow` (1–100). Query `chromatogram` for at most 100 original-orientation called bases and sample pages up to 500 rows. Base coordinates are 1-based inclusive; trace sample indices are explicitly zero-based. Preserve the returned quality encoding.
- `set_read_pileup_options`, `select_read`, and `clear_read_selection`: use the same MAPQ/flag/strand filters, sort, base/soft-clip visibility and selected-read state as the UI. Query `reads` or `coverage` with `reference`, `start`, and `end` plus optional pagination, without `record` or `trackId`. Query `read-pileup-state` with only `sessionId` and `target`. For `read-detail` and `select_read`, use exact `trackId` and `sourceReadIndex` from a current reads query; that zero-based index belongs to the loaded track array, not the source-file position.
- `set_quality_view_options`: expand quality distributions, methods or named data tables. Query `quality-report` for the same bounded report shown in the viewer and `sequence-ui-state` for display options; a pending job is not a completed report.
- `clear_sequence_selection`: clear the active sequence range and pinned feature.
- `focus_alignment_cell`, `focus_alignment_reference_coordinate`, `select_alignment_columns`, or `select_alignment_rows`: focus an alignment row/column, map an ungapped active-reference coordinate into the alignment, or select columns/rows. Columns are 1-based inclusive; an empty `rows` array clears row selection.
- `set_alignment_reference`: use `consensus`, `none`, or an exact row ID or label.
- `search_alignment` and `navigate_alignment_search_hit`: search the alignment motif view and move through the resulting hits.
- `filter_alignment_rows`, `set_alignment_row_visibility`, or `show_all_alignment_rows`: filter rows by label or description, hide or show exact row IDs or labels, or restore every row.
- `set_alignment_view_options`: set molecule interpretation, analysis and search scopes, cell width, color mode, residue palette, annotation and metric tracks, row sort key/direction, sequence-logo help, identical-as-dots rendering, and RNA structure overlays. Use modality-compatible colors and palettes.
- `clear_alignment_selection` or `reset_alignment_view`: clear the current focus/selection or restore default alignment controls.
- `compute_alignment_guide_tree`: compute the legacy exploratory UPGMA view. Prefer `sequence.run_analysis` with `build-tree` for the graphical NJ/UPGMA workbench and provenance-bearing result.

Choose settings compatible with the current record and its available tracks. After the acknowledged controls, read `sequence.get_context` to verify the resulting display and selection on the same viewer, then briefly say what changed. Do not claim that a requested state changed when the tool reports `applied: false`.

## Scientific Workbench Workflow

Treat immediate view controls and scientific operations differently. Controls are synchronous and idempotent. Analyses, alignment, edits, evidence loading, exports, and sessions are typed workbench operations tied to the current `viewerSessionId`.

- Use `sequence.query_viewer` whenever context says a collection is truncated or the requested record, feature, sequence/quality range, alignment row/column, metric, hit, track, tree node, variant, coverage bin, read, job, or artifact is not in the compact context. Follow `nextCursor`; never guess omitted state.
- Use `sequence.run_analysis` for statistics, genetic-code-aware translation, ORFs, restriction sites/digest, exploratory primer candidates, distance matrices, graphical NJ/UPGMA trees, or document-scoped `quality-report`. The quality report accepts an optional `adapterSequence` of 8–64 A/C/G/T bases; it samples retained reads rather than the current selected range. Select the relevant range/rows for other analyses when appropriate. Report the returned engine, parameters, limitations, truncation, and coordinate system.
- Use `sequence.align` for a new pairwise or center-star alignment from exact record/row IDs. It produces a derived aligned-FASTA artifact and never overwrites the source. The built-in engine is exploratory; hand off publication or large-production MSA work to an approved MAFFT/Clustal workflow and reopen the returned alignment.
- Use `sequence.edit_copy` for typed sequence insert/delete/replace/reverse-complement/rotation operations or alignment gap/column/row/group/sort operations. Edits apply only to the in-memory copy. Undo/redo restore copy history and accept only `sessionId` and `operation`, without `record` or other edit fields. Export the copy explicitly.
- Use `sequence.manage_annotations` to add, update, delete, or import features from a loaded annotation track. Preserve qualifier and coordinate provenance. Resolve duplicate IDs explicitly.
- Use `sequence.load_track` for GFF3/GTF/BED, VCF, SAM, or indexed regional BAM/CRAM evidence on an opened sequence/reference. BAM/CRAM are not primary viewer documents. BAM needs a BAI/CSI index (default: `<path>.bai`). CRAM needs a CRAI index (default: `<path>.crai`) and usually a matching `referencePath` FASTA. Omit `referencePath` for BAM and all other formats. Keep indexed windows at or below 100,000 bases. Inspect mapping diagnostics, loaded-source coverage and disclosed downsampling before interpreting evidence; base-level pileup detail requires a window of at most 200 bases. Editing the reference copy invalidates imported evidence placement rather than remapping it; undo the edit or reopen a matching source reference. Display-only reverse-complement projection does not invalidate the source binding.
- Use `sequence.export_artifact` for sequence/alignment data, annotations, tables, SVG figures, Newick, or JSON. Private artifact persistence remains the default. When an independently authenticated plugin workspace capability exists and the user explicitly asks to publish beside the opened source, pass `destination: {kind: "workspace", base: "opened-source", relativePath: "..."}` with a source-relative path; safe `..` is allowed only when the server confirms the result remains inside the same active workspace root. Never pass an absolute, drive, UNC, URL, or guessed path. In-view **Publish** actions use an explicit, source-bound workspace Save As confirmation. If that capability is unavailable, plugin publication must fail closed. Codex may separately use its own authorized workspace tools to save a small, exact, user-requested result that is already legitimately available in model-visible analysis output, such as the starter's completed Newick tree; identify Codex as publisher and do not claim the plugin published or signed it. Otherwise preserve the private artifact and explain the missing authorization; do not offer browser **Download**, which Codex blocks. If clipboard access is unavailable, in-view Copy actions expose at most 128 KiB of exact, read-only selectable text for manual **Command+C** or **Ctrl+C**; larger selections require authorized workspace publication. Do not call the app-only preflight or chunk tools directly. Use `sequence.save_session` and `sequence.restore_session` for round-trippable view, selection, copy, jobs, tracks, and derived artifacts; sessions are never workspace-published. Use `sequence.cancel_job` with the exact running job ID.

Prefer contextual workflows over listing every capability. Examples: select a CDS then translate; select a target then design primers; load a VCF then query variants in the visible region; select alignment rows then realign/build a tree/export Newick. The viewer keeps its CSP closed: use approved Codex research or science tools for BLAST, structure linkage, RNA folding, CRISPR/cloning, genome-wide primer specificity, publication-grade phylogenetics, or other external engines, then load or open their traceable artifacts in the viewer.

## Supported Artifacts

- Sequence-first files: FASTA, FASTQ, GenBank, EMBL, and common plain biological sequence suffixes.
- Bounded binary imports: original-orientation processed ABIF 1.01 (`.ab1`, `.abi`), SCF 2.00/3.00 (`.scf`), and verified-revision SnapGene DNA (`.dna`). Unsupported versions/revisions fail closed; these are not full native application document imports.
- Alignment-first files: aligned FASTA, AFA, CLUSTAL, Stockholm, A2M, A3M, MSF, PHYLIP, NEXUS, and PIR.
- Ambiguous rectangular FASTA files open in the combined viewer with an in-view Sequence / Alignment mode choice.

## Boundaries

- Do not read, inline, truncate, or summarize file contents merely to open the viewer.
- Do not pass a workspace path, `file://` URI, or fabricated `codex-resource://` URI to `sequence.open`.
- When the MCP host exposes active local roots, `sequence.open_from_chat` confines both relative and absolute paths to those roots. When roots are unavailable, it requires an exact absolute local path.
- `sequence.open_from_chat` accepts only supported files.
- Do not infer the user's current selection from the source file after the app is open. Use the viewer context because the user may have changed modes, rows, columns, or selections interactively.
- Treat FASTQ quality as Phred+33 unless the source explicitly establishes another encoding. Detailed quality reports use the first eligible retained reads, at most 1,000 reads and 100,000 bases; report their actual scope separately from streaming summaries. They are descriptive reports, not FastQC runs, automatic contamination/pass/fail assessments or whole-file QC.
- Binary inputs are capped at 8 MiB and saved sessions at 512 KiB. Do not drop trace signals silently to fit a session or claim native ABIF/SCF/SnapGene export. Preserve originals for data not represented by the import. SCF confidence bytes are source confidence, not assumed calibrated Phred scores; trace views/queries retain original read orientation.
- Treat built-in NJ/UPGMA trees, pairwise/center-star alignment, and primer scoring as exploratory, not publication-grade or experimentally validated inference.
- Never silently choose genetic code 1 when source features declare a different `transl_table` or multiple tables. Use the source-derived code when unique; otherwise specify or ask for `geneticCodeId`.
- Treat `activeTarget` as authoritative. Mention `transientFocus` separately only when both are populated and refer to different coordinates/rows; pointer movement must not replace the user's selected target.
- Preserve compound GenBank and EMBL feature segments; do not describe the gaps between joined segments as part of the feature.

## Follow-Up Questions

When answering from the mounted viewer context, prioritize the literal current state before adding biological interpretation. Useful context includes `activeTarget`, distinct `transientFocus`, coordinate semantics, source identity, current mode, record/feature inventories, selected sequence segments, genetic code, FASTQ QC, exact compound features, selected alignment rows/columns, reference mappings, conservation, tree provenance, loaded evidence/downsampling, jobs, and derived artifacts. Query omitted pages instead of inferring them from source bytes.

SHA-256: 49a56c831b0c767533a133ecc67c541b49c27a3753c6605ee24f03e2c43ac4f9