← Files Biological Sequence & Alignment ViewerARCHIVED FILE

CAPABILITY_MATRIX.md

16.4 KB · Sep 30, 2026 · 23:01 UTC

↓ Download file

# Sequence Viewer Capability Matrix

This is the capability boundary of the source implementation, not a release qualification or a claim of complete Jalview, IGV, SnapGene, Benchling, or FastQC equivalence. The viewer combines sequence, alignment, trace, quality, and regional read inspection in one workspace; specialist analysis and native application document fidelity remain separate concerns.

## Implemented inspection surfaces

| Area | Implemented | Boundary |
| --- | --- | --- |
| Sequences and annotated constructs | FASTA/FASTQ, GenBank/EMBL; record browser; motif search; translation and ORFs; linear, circular, and split maps; directional, selectable annotations and compact linked labels; compound and origin-spanning locations; reverse-complement projection | Source coordinates remain 1-based inclusive. Maps prioritize the selected feature within 200 annotations, six collision-separated lanes, and 1,200 processed segments. Up to 12 names remain visible when the searchable 30-row annotation index is closed; omitted geometry and minimum-width markers are disclosed. Remote intervals are not drawn on the local molecule. |
| Multiple-sequence alignments | Aligned FASTA, CLUSTAL, Stockholm, A2M/A3M, MSF, PHYLIP, NEXUS, PIR; virtualized grid; references, differences, conservation, metric tracks, logos, annotations, row/column selection, sorting/grouping, safe edits, distance matrices and exploratory NJ/UPGMA trees | Built-in pairwise/center-star alignment and guide trees are exploratory. No bundled MAFFT, full phylogenetic inference, or full Jalview analysis ecosystem. |
| Regional mapped reads | SAM and indexed BAM/CRAM evidence; CIGAR-aware aligned blocks, insertions, deletions, reference skips, soft clips, mismatches and base qualities; coverage; MAPQ/flag/strand filters; sorting; selected-read details and an unambiguous displayed mate | Open a sequence/reference first and attach the read track. BAM/CRAM are not primary genome-browser documents. No aggregate splice-junction arcs, paired-fragment packing, quality-weighted allele fractions, arbitrary tag coloring, or full IGV/JBrowse track ecosystem. |
| Detailed FASTQ quality | Per-cycle mean/min/max quality and composition, read length/GC/mean-quality distributions, exact repeated reads, frequent 7-mers, and an optional supplied adapter screen; plots and data tables use the same bounded report | A descriptive subset report, not a FastQC execution, equivalence claim, contamination diagnosis, pass/fail assessment, or library-complexity estimate. Bounds and sampling are described below. |
| Sanger traces | ABIF `.ab1`/`.abi` and SCF `.scf`; four-channel traces, called bases, peak positions, available confidence values, source-coordinate selection, zoom/window navigation, and paginated signal queries | ABIF 1.01 processed `DATA9–12` in original source orientation only. SCF 2.00/3.00 with 1- or 2-byte samples and supported code sets 0/2/3/4. No re-basecalling, trace editing, native trace export, or claim to reproduce instrument/raw-data analysis. |
| SnapGene import | `.dna` DNA sequence, topology, supported features, compound locations, qualifiers, notes and primer sites; unsupported properties are reported | Verified export/import revision pairs only: `13:11`, `13:12`, `14:16`, `15:19`, `15:20`. Not a complete SnapGene document import: unsupported packets, history, application settings and chemistry remain in the original. No native `.dna` export. |
| Workbench and outputs | Reversible copy edits; annotation and evidence management; bounded analyses/jobs; typed exports; private sessions; authorized create-new workspace publication with provenance | Source files are unchanged by default. Workspace writes retain their authorization and confirmation gates. No ELN registry, shared project collaboration, wet-lab execution, or comprehensive cloning/assembly suite. |

The compact toolbar and tool panels keep the sequence or alignment as the primary surface. Tool panels, nested disclosures, record lists, annotation lists, read details, trace windows, and quality tables can be addressed through typed operations rather than browser automation.

## UI-to-agent mapping

Use the mounted viewer's `viewerSessionId` as `sessionId`. Control actions below belong to `sequence.control_viewer`; query targets belong to `sequence.query_viewer`. An accepted request or started job is not a completed result: inspect `applied`, returned state, job status, and current viewer context.

| UI action | Agent operation and readback |
| --- | --- |
| Show/hide toolbar; move between chat and side pane; switch Sequence/Alignment | `set_toolbar_visibility`, `set_display_mode`, `set_mode`. `fullscreen` is the side-pane host mode; `inline` returns to chat. |
| Open or close a tool panel | Query `workbench-panels`; call `set_workbench_panel` with its exact `group` and `panel` ID, or `panel: null` to close. Groups are `sequence-display`, `sequence-tools`, and `alignment-tools`. Discover the available panels for the current artifact rather than assuming every panel is mounted. |
| Expand/collapse a nested section | Query `workbench-disclosures` with optional `mode`, `limit` up to 100, and `cursor`; call `set_workbench_disclosure` with the returned `disclosureId` and `expanded`. Expanding reveals its containing panel and registered ancestor sections. Unknown, unavailable, or guarded sections return `applied: false`. |
| Dismiss registered copy feedback or a session error | Query `workbench-feedback` with optional `mode`, `limit` up to 100, and `cursor`; call `dismiss_workbench_feedback` with the exact returned `feedbackId`. Only registered copy/session-error handlers are available. This cannot approve, confirm, cancel or overwrite a source write; feedback queries do not expose the manual-copy sequence. |
| Search, sort, page, or expand the record browser; choose a record | `set_sequence_record_browser` with `query`, `sortBy`, zero-based `page`, or `expanded`; `set_sequence_record` chooses an exact/unambiguous record. Query `records` for paginated metadata and `sequence-ui-state` for actual browser state. |
| Change sequence map, orientation, display tracks, palette, wrapping, or genetic code | `set_sequence_view_options` with `layout`, `orientation`, `showFeatures`, `showQuality`, `showTranslation`, `palette`, `wrapWidth`, `geneticCodeId`, or `synchronizedViews`; query `sequence-ui-state`. |
| Choose a soft, Monochrome, or scientific residue palette | `set_sequence_view_options.palette` or `set_alignment_view_options.residuePalette`: `muted-nucleic-acid` is the DNA/RNA default, `muted-amino-acid` the protein default, and `neutral` selects Monochrome. Query `sequence-ui-state.palette` for active `id`, `defaultId`, compatible `options`, and any `restorationWarning`. Alignment context exposes `display.residuePalette`, `defaultResiduePalette`, and `compatibleResiduePalettes`. Compatible saved choices and canonical scientific palette options are preserved. |
| Search/page/expand the map annotation index; select an annotation | `set_sequence_annotation_index` with `query`, zero-based `page`, or `expanded`; `select_sequence_feature` or `navigate_sequence_feature`. Query `features` for exact IDs, qualifiers and segment coordinates, including annotations outside the drawn overview. |
| Navigate, search, select bases, or select across the origin | `focus_sequence_coordinate`, `search_sequence`, `navigate_sequence_search_hit`, `select_sequence_range`, `clear_sequence_selection`. Origin selection uses `wraparound: true` and source `start > end`; `set_sequence_view_options.originRangeExpanded` controls the form. Query `sequence-range`, `quality`, or `search-hits` for bounded scientific data. |
| Move/zoom a chromatogram; select called bases | `set_chromatogram_view_options` with 1-based `firstBase` and `basesPerWindow` from 1–100; use `select_sequence_range` for source base selection. Query `chromatogram` with `start`/`end` covering at most 100 called bases and sample pages up to 500 rows. `sequence-ui-state` reports the trace view settings. |
| Filter/sort read pileups; show bases or clips | `set_read_pileup_options`: `minimumMappingQuality`, `strand`, `includeDuplicates`, `includeQcFailed`, `includeSecondary`, `includeSupplementary`, `includeUnknownMappingQuality`, `sortBy`, `showAllBases`, `showSoftClips`. Query `read-pileup-state`, `reads`, and `coverage`. |
| Inspect or clear a selected read; inspect its available mate | `select_read` / `clear_read_selection`; query `read-detail`. Use exact `trackId` and zero-based `sourceReadIndex` from a current `reads` result. This index identifies the materialized track array, not a source-file ordinal or byte offset. Mate links apply only when an unambiguous mate is available in the displayed sample. |
| Run quality analysis or change its adapter screen | `sequence.run_analysis` with `analysis: "quality-report"` and optional `adapterSequence` of 8–64 A/C/G/T bases; query `jobs`, then `quality-report`. This is document-scoped, not a selected-range analysis. |
| Open quality distributions, methods, or data tables | `set_quality_view_options` with `distributionsExpanded`, `methodsExpanded`, or `expandedTables`. Table IDs are `cycle-quality`, `cycle-composition`, `read-length`, `read-gc`, `read-mean-quality`, `repeated-sequences`, and `frequent-kmers`. Query `sequence-ui-state` and `quality-report`. |
| Change alignment appearance, metric tracks, scopes, or row ordering | `set_alignment_view_options`, including `enabledMetricTracks`, `showSequenceLogoHelp`, `rowSortKey`, `rowSortDirection`, color/palette and analysis/search scopes. Query `metrics`, `annotations`, and `rows`; current options also appear in viewer context. |
| Select/filter/hide alignment rows or columns; choose reference; navigate matches | `select_alignment_rows` (empty `rows` clears), `select_alignment_columns`, `filter_alignment_rows`, `set_alignment_row_visibility`, `show_all_alignment_rows`, `set_alignment_reference`, `focus_alignment_cell`, `focus_alignment_reference_coordinate`, `search_alignment`, and `navigate_alignment_search_hit`. Query `rows`, `columns`, `search-hits`, and `tree-nodes` as needed. |
| Analyze, realign, edit, annotate, attach evidence, export, save/restore, or cancel | `sequence.run_analysis`, `sequence.align`, `sequence.edit_copy`, `sequence.manage_annotations`, `sequence.load_track`, `sequence.export_artifact`, `sequence.save_session`, `sequence.restore_session`, `sequence.cancel_job`. Query `jobs`, `artifacts`, and `tracks` for status and provenance. These operations preserve the same validation and authorization boundaries as the UI. |

Registered sections can remain discoverable while visually hidden. Shared panel/disclosure handlers reject closes that would hide protected pending approval or manual-copy state. Registered copy/session-error feedback can be explicitly dismissed through its own action; that is not approval of a user's confirmation. Paginated queries expose more than the current rendered page, but still only within the loaded-source, query, and completion budgets. Follow `nextCursor` instead of assuming omitted content.

## Scientific and resource limits

- **Binary import:** at most 8 MiB per ABIF/SCF/SnapGene input. Chromatograms additionally cap 200,000 called bases, 500,000 samples per channel, and 10,000 ABIF directory entries. Malformed offsets, unsupported versions and invalid peak/channel relationships fail closed. A successfully opened trace does not necessarily fit the independent 512 KiB saved-session budget; signals must not be silently dropped to make a session fit.
- **Trace meaning:** ABIF processed signals and matched base-call/peak/confidence revisions are not raw instrument channels. Reverse-complemented ABIF sources are rejected. The trace view and its signal queries retain original read orientation even when the sequence display is reverse-complemented. SCF confidence bytes are retained as source confidence, not relabeled calibrated Phred scores; ambiguous calls have no invented called-base confidence. SCF clipping hints are metadata, not automatic trimming.
- **FASTQ sampling:** the detailed report considers at most 5,000 retained reads and analyzes the first 1,000 eligible complete quality-bearing reads, at most 100,000 bases, skipping reads over 10,000 bases. It is source-order sampling, not random sampling or whole-file QC. Up to 80 cycle bins, 40 histogram bins, six repeated sequences and six frequent patterns fit a 48 KiB structured report. The separate streaming FASTQ summary has its own larger scope; never equate its population with the detailed report's analyzed subset.
- **FASTQ methods:** Phred+33 is assumed. Cycle plots show mean/minimum/maximum, not quartiles. Read GC excludes ambiguous bases from its denominator. Repeated reads are exact full forward sequences; frequent 7-mers are observations, not enrichment tests. Adapter screening uses only the supplied contiguous forward sequence, without mismatches, partial matches, reverse-complement screening or automatic contamination identification.
- **Read evidence:** indexed loads require a BAI/CSI for BAM or CRAI for CRAM and an appropriate reference when needed; the regional load limit is 100,000 bases. The UI draws at most 100 reads and shows individual bases/qualities only in windows of 200 bases or fewer. Coverage uses filtered, loaded alignments before display sampling, counts `M`, `=` and `X`, and excludes deletions/reference skips. Source truncation, unavailable/invalid CIGAR and processing limits make coverage explicitly incomplete; it is never extrapolated to unseen reads. MAPQ 255 means unavailable, not high confidence.
- **Edited references:** editing the reference sequence copy invalidates imported evidence coordinates. Tracks retain their original coordinates, but read placement, coverage, variants and imported annotations become unavailable on the edited copy. Undo the edit or reopen a matching source reference; no implicit remapping occurs. A display-only reverse-complement projection is not such an edit.
- **Outputs:** supported text, table, figure and session exports are derived artifacts, not binary/native ABIF, SCF, SnapGene, BAM or CRAM round trips. Sessions remain capped at 512 KiB and private artifacts at 8 MiB. Native application history, hidden annotations/settings and original binary fidelity require retaining the original file. External engines can produce new traceable artifacts to reopen, but the viewer does not silently run external tools.

## Evidence and remaining qualification

Source tests exercise geometry, coordinate/orientation semantics, malformed binary input, QC bounds, CIGAR behavior, control/query schemas and UI handlers. Generated conformance fixtures are tests, not customer data or evidence of biological validity.

Development binary checks use three real Biopython ABIF examples and two BioJava SCF examples. Reference checks include Biopython expected sequences/qualities and the complete called sequence, peaks and channels of `A6_1-DB3.ab1`. A separate reproducible direct byte/spec decoder additionally matched all 230,520 signal sample values across all four channels in those five files, along with called bases, peak offsets, available ABIF qualities and SCF candidate confidences. The SCF test decoder uses separate modular prefix-sum passes. This is a direct byte/spec comparison, not execution of a third-party decoder or native application. SnapGene receipts cover eight real Biopython fixtures, selected golden length/topology/feature assertions, compound/origin paths and qualifier normalization. They do **not** establish every annotation/qualifier, full independent sequence equivalence, executed Biopython differential comparison, or native SnapGene UI equivalence.

The development receipts are `sequence-binary-reference-validation.json`, `sequence-binary-direct-decoder-independent-review.json`, and `sequence-snapgene-reference-validation.json`; the direct comparison has a matching `.mjs` test script and pins source/fixture digests. These local receipts and temporary reference binaries are not shipped. The parser was independently authored with reference to format documentation and Biopython sources/tests, not under a clean-room process. Local/source tests and these receipts do not establish installed-host adoption, marketplace publication, universal file compatibility, or full specialist-product parity.

Source-checkout maintainers can inspect `src/viewer-commands.ts`, `src/viewer-operations.ts`, `src/viewer-query.ts`, `src/sequence/use-sequence-interface-commands.ts`, `src/ui/workbench-tool-controller.ts`, and the focused map, trace, QC and read-pileup tests. Current release/host qualification remains a separate gate from this matrix.

SHA-256: f87a16eb2a2271869e6aa23c7eed9f8c2d35c8925b14dec833ee7b0848892db7