← Plugin catalog
Education & Research

Life Sciences NGS Analysis

OpenAI v1.0.3

Publisher description

From the marketplace listing

A guided intake, routing, and execution plugin for next-generation sequencing workflows. It helps Codex inspect local sequencing inputs, ask only the missing assay-specific questions, choose public or freely accessible runtime-installable packages where possible, check existing tool availability before any install, and execute supported local workflows with validation, logs, manifests, QC reports, and artifact indexes. It includes deeper decision skills for BCL demultiplexing, FASTQ QC execution and interpretation, germline, somatic and UMI-panel DNA variants, bulk RNA-seq count generation and differential expression, ATAC-seq, ChIP-seq/CUT&RUN/CUT&Tag, and embedded post-count scRNA-seq QC.

Language: English · Automatically detected from descriptions.

Files & skills

File archives

Plugin package81 files · 283 KBBrowse files →
Skill instructions
ngs-amplicon-microbiome5.66 KB

View saved version →

---
name: ngs-amplicon-microbiome
description: Kick off public 16S, 18S, ITS, COI, or other marker-gene amplicon microbiome workflows using nf-core/ampliseq, QIIME2, DADA2, and Cutadapt.
---

# Amplicon Microbiome

Use this skill for marker-gene microbiome analysis from amplicon FASTQs.

## Essential Inputs

Confirm:

- marker region: 16S, 18S, ITS, COI, or custom
- primer sequences and orientation
- paired-end or single-end reads
- whether reads should be merged
- taxonomy database and version
- sample metadata
- endpoint: ASV table, taxonomy, diversity, differential abundance, or plots

## Public Defaults

Prefer `nf-core/ampliseq` for reproducible end-to-end runs. Use QIIME2 or DADA2 directly when the user wants notebook-level control or an existing lab protocol requires it.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline amplicon_microbiome --emit-install-plan
```

## Local Execution Package

For FASTQ intake/QC before primer, ASV, and taxonomy decisions, use:

```bash
python plugins/ngs-analysis/scripts/run_fastq_assay_package.py \
  --lane amplicon_microbiome \
  --sample-sheet amplicon_samples.tsv \
  --execute
```

This validates read paths and structure, runs seqkit stats and FastQC/MultiQC when available, and writes `amplicon_analysis_status.json`. The runner now also emits `methods/amplicon_methods.json` plus a concrete backend handoff bundle under `workflow/` so primer, denoiser, truncation, normalization, and taxonomy choices are machine-readable even before a full backend is run.

If the user asks for a full amplicon analysis rather than QC/readiness, do not treat FASTQs alone as sufficient. Require primer sequences, primer orientation, taxonomy database plus version, and sample metadata before presenting the run as analysis-ready. Without that context, run the local execution package and describe the result as a read-QC/readiness bundle only.

For backend ASV/taxonomy/diversity execution when primers, metadata, and taxonomy resources are available, use:

```bash
python plugins/ngs-analysis/scripts/run_amplicon_microbiome.py \
  --sample-sheet amplicon_samples.tsv \
  --backend qiime2 \
  --primer-forward GTGYCAGCMGCCGCGGTAA \
  --primer-reverse GGACTACNVGGGTWTCTAAT \
  --taxonomy-classifier silva-138-classifier.qza \
  --metadata sample_metadata.tsv \
  --execute
```

Use `--backend dada2` for a direct R/Bioconductor ASV path. The plugin includes `workflows/amplicon_microbiome/run_dada2_backend.R`; the runner checks for `Rscript` and the `dada2` R package before execution, then writes normalized ASV, representative-sequence, read-retention, and optional taxonomy tables under `tables/`.

For nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline ampliseq`.

The direct backend runner also emits `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. The resource check is advisory by default when a QIIME classifier is supplied directly; add `--bundle-root silva_138_amplicon=<path>`, `--include-optional-resources`, and `--require-resource-plan` when missing registered taxonomy databases should block readiness.

The backend runner writes native normalized tables when QIIME2/DADA2/nf-core outputs are present:

- `tables/asv_table.tsv`
- `tables/representative_sequences.fasta` for direct DADA2 runs
- `tables/taxonomy.tsv`
- `tables/read_retention.tsv`
- `tables/amplicon_backend_summary.json`
- `tables/alpha_diversity.tsv`, `tables/bray_curtis_distance.tsv`, and `tables/top_taxa_or_features.tsv` when a normalized ASV/feature table is available

QIIME2 BIOM-only feature-table exports are recorded as requiring conversion, with a `biom convert` command in the backend summary. Do not claim diversity or taxonomy interpretation unless these normalized tables or equivalent supplied inputs exist.

## Kickoff Pattern

nf-core preflight run:

```bash
nextflow run nf-core/ampliseq \
  -profile test,docker \
  --outdir results/ampliseq_test
```

Before a real run, verify primer trimming and truncation choices from read-quality profiles.

## Visualization Outputs

The local FASTQ package always writes `visualizations/index.html` and `visualizations/visualization_manifest.json`. With only FASTQs, this is a read-QC/readiness bundle. If an ASV/feature table is available, pass it to the runner with `--asv-table` to generate alpha diversity, Bray-Curtis PCoA, and rarefaction artifacts. If a feature taxonomy table is available, pass `--taxonomy-table` to generate taxa barplots. When downstream tables are labeled synthetic or contain sample columns that are not present in the real sample sheet, the runner marks the run review-only and blocks beta-diversity/PCoA unless `--allow-synthetic-diversity` is set explicitly.

The run also emits `qc_verdict.json` and, for amplicon runs, `qc_interpretation.json` with machine-readable reason codes, a readiness verdict, and follow-on command templates for generating ASV/taxonomy tables and re-rendering plugin-native plots. Backend runs additionally write `tables/amplicon_backend_summary.json` so exported ASV, taxonomy, read-retention, and BIOM-conversion status are auditable. When a normalized ASV/feature table is available, the backend runner also writes `tables/amplicon_diversity_summary.json`, `visualizations/amplicon_backend_dashboard.html`, and SVG plots for sample depth, Shannon diversity, and top taxa/features. If the ASV table is absent, these outputs remain explicitly unavailable rather than inferred from FASTQ QC.

## Guardrails

- Do not choose truncation lengths before looking at quality distributions.
- Do not mix taxonomy database versions without recording them.
- Preserve negative controls and extraction blanks in metadata.

Referenced files: 1

ngs-analysis-router4.31 KB

View saved version →

---
name: ngs-analysis-router
description: Route BCL, FASTQ, BAM/CRAM, count-matrix, or VCF sequencing requests to the right public NGS analysis skill and ask only the missing assay-specific setup questions.
---

# Life Sciences NGS Analysis Router

Use this skill as the top-level entrypoint for ambiguous or broad sequencing-analysis requests.

## Start Here

Inspect the available inputs before asking the user questions. Look for:

- Illumina run-folder files: `RunInfo.xml`, `RunParameters.xml`, `SampleSheet.csv`, `Data/Intensities/BaseCalls`
- FASTQs: `*.fastq`, `*.fq`, `*.fastq.gz`, `*.fq.gz`
- BAM/CRAM/VCF: `*.bam`, `*.cram`, `*.vcf`, `*.vcf.gz`
- count matrices: `matrix.mtx`, `features.tsv`, `barcodes.tsv`, `*.h5`, `*.h5ad`, `*.rds`
- metadata: sample sheets, design files, target BEDs, reference FASTA/GTF, primer files

Read `references/intake-schema.json` and `references/pipeline-registry.json` when forming the route.

## Intake Rules

Ask the smallest set of missing questions needed to choose a defensible pipeline. Do not ask the full questionnaire if file inspection already answers a field.

Always resolve:

- input type
- assay type
- desired output
- organism/reference
- paired-end vs single-end when FASTQs are involved
- any assay-specific design file or metadata required for the requested result
- runtime constraints: local/HPC/cloud, container availability, and whether installs are allowed

For human data, ask whether cloud upload is allowed before suggesting BaseSpace, Terra, DNAnexus, or any cloud path.

## Routing

Route to one leaf skill:

- BCL run folder or demultiplexing: `ngs-bcl-to-fastq`
- QC/trimming only: `ngs-fastq-qc`
- WGS/WES/panel variants: `ngs-dna-variant-calling`, then a subtype skill when the analysis model is clear
- germline WGS/WES/panel variants: `ngs-dna-germline-variants`
- tumor-normal or tumor-only somatic variants: `ngs-dna-somatic-variants`
- UMI, duplex, or low-frequency targeted panels: `ngs-dna-umi-panel-variants`
- bulk RNA-seq kickoff: `ngs-bulk-rnaseq`
- bulk RNA-seq FASTQ-to-count QC: `ngs-bulk-rnaseq-counts-qc`
- bulk RNA-seq differential expression from counts: `ngs-bulk-rnaseq-differential-expression`
- single-cell or single-nucleus FASTQ-to-matrix kickoff: `ngs-scrna-seq`
- single-cell or single-nucleus post-count QC/annotation/UMAP: `scrna-seq-qc`
- epigenomics kickoff: `ngs-epigenomics-peaks`
- ATAC-seq QC/peaks/accessibility: `ngs-atacseq-peaks-qc`
- ChIP-seq, CUT&RUN, or CUT&Tag QC/peaks: `ngs-chip-cutrun-peaks-qc`
- 16S/18S/ITS/COI amplicons: `ngs-amplicon-microbiome`
- shotgun metagenomics: `ngs-shotgun-metagenomics`
- runtime/package setup only: `ngs-runtime-env`

Prefer public, runtime-installable packages and nf-core workflows. Surface license/EULA/account boundaries before using proprietary or cloud tools.

## Preflight

Before proposing installation or execution, run a preflight plan from the repo root:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline <pipeline_key> --emit-install-plan
```

When the user needs an approval-ready install handoff, write persistent install artifacts:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline <pipeline_key> --manager micromamba --install-plan-outdir runtime_readiness/<pipeline_key>_install
```

Treat `install_plan.json` as the canonical review artifact. `install_commands.sh` is generated from the same plan and stays review-only unless the user explicitly approves execution with `NGS_RUN_INSTALL_COMMANDS=1`.

For reference- or database-heavy pipelines, also create a resource plan before saying the workflow is runnable:

```bash
python plugins/ngs-analysis/scripts/ngs_reference_manager.py plan --pipeline <pipeline_key> --genome-build <build> --outdir resource_readiness/<pipeline_key>
```

Use `--include-optional` for shotgun, amplicon, or motif-enabled epigenomics runs when optional databases materially affect the requested output.

Use `--network-checks` only when the user allows network checks. Use `--install-missing --yes` only when the user explicitly asks to install.

## Output Contract

Return:

1. the routed analysis type and confidence
2. missing essential parameters, if any
3. recommended public pipeline or package family
4. local tool preflight summary
5. preflight-first command or next concrete action
6. caveats around licenses, cloud upload, database size, and reference data

Referenced files: 1

ngs-atacseq-peaks-qc3.45 KB

View saved version →

---
name: ngs-atacseq-peaks-qc
description: Run or plan ATAC-seq QC, alignment, TSS enrichment, fragment-size, blacklist, peak-calling, consensus peak, and differential accessibility workflows.
---

# ATAC-seq Peaks QC

Use this skill for ATAC-seq accessibility analysis from FASTQ or BAM. If the assay is ChIP-seq, CUT&RUN, CUT&Tag, or antibody-targeted enrichment, use `ngs-chip-cutrun-peaks-qc`.

## Essential Inputs

Confirm:

- FASTQ/BAM inputs and paired-end status
- organism, genome build, blacklist, and mitochondrial contig names
- biological replicates, conditions, batches, and sample metadata
- whether the target is QC only, peaks, consensus peaks, bigWigs, or differential accessibility
- whether Tn5 shifting is handled by the chosen workflow
- desired peak caller and downstream matrix generation

## Route

Prefer `nf-core/atacseq` for full reproducible processing. Use direct MACS2 only when BAMs are already aligned, duplicate/blacklist handling is known, and the user wants focused peak calling.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline atacseq_peaks_qc --emit-install-plan
```

For compact read-level intake/QC, use the shared epigenomics execution package:

```bash
python plugins/ngs-analysis/scripts/run_fastq_assay_package.py \
  --lane epigenomics_peaks \
  --sample-sheet atac_samples.csv \
  --execute
```

For local-light ATAC alignment, peaks, FRiP, TSS, bigWig tracks, and consensus peaks from FASTQ or prepared BAMs, use the dedicated ATAC runner:

```bash
python plugins/ngs-analysis/scripts/run_atacseq_peaks_qc.py \
  --sample-sheet atac_samples.csv \
  --bowtie2-index /refs/GRCh38/bowtie2/genome \
  --genome-size hs \
  --blacklist-bed /refs/GRCh38/blacklists/encode_blacklist.bed \
  --tss-bed /refs/GRCh38/tss.bed \
  --execute
```

This runner emits `qc/atacseq_qc_summary.{tsv,json}`, `qc/atacseq_qc_dashboard.html`, native SVG FRiP/peak and insert-size plots, browser-track handoff files under `tracks/`, and TSS profile/heatmap commands when `--tss-bed` is supplied. Add `--run-motifs --motif-genome <genome>` when HOMER motif enrichment should be part of the backend run.

It also emits `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. The resource check is advisory by default for local-light runs; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when missing registered reference bundles should block readiness.

For nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline atacseq`.

## QC Gates

Review before biological interpretation:

- read depth, alignment rate, duplicate rate, and mitochondrial fraction
- insert-size periodicity/nucleosome pattern
- TSS enrichment and FRiP score when available
- blacklist overlap and peak count per sample
- replicate concordance and consensus peak support

Do not proceed to differential accessibility if replicate quality or metadata is insufficient.

## Outputs

Produce:

- sample sheet and workflow command/profile
- QC summary and failed-sample flags
- narrowPeak/BED peak sets, consensus peaks, bigWigs, browser-track manifests, browser-track preview HTML, native QC dashboard/SVG plots, TSS plots, and peak-count matrix when requested
- motif summary files when a motif backend is requested
- differential-accessibility design and contrasts if applicable
- caveats for low TSS enrichment, high mitochondrial reads, weak replicate concordance, or poor FRiP

Referenced files: 1

ngs-bcl-to-fastq3.85 KB

View saved version →

---
name: ngs-bcl-to-fastq
description: Validate Illumina BCL run folders and sample sheets, plan demultiplexing, review index/UMI/lane choices, run BCL-to-FASTQ conversion, and interpret demux metrics while surfacing license/download boundaries.
---

# BCL To FASTQ

Use this skill when the input is an Illumina BCL run folder or the user asks to demultiplex a sequencing run. This is a deep demultiplexing and run-validation skill, not only a command wrapper.

## Essential Inputs

Confirm:

- run folder path with `RunInfo.xml`
- sample sheet path and format
- output directory
- instrument/run metadata from `RunInfo.xml` and `RunParameters.xml`
- lane handling: split by lane or combine lanes
- index mismatch tolerance
- index read structure and dual-index orientation
- UMI layout, if any
- whether adapter trimming/masking should happen during conversion
- whether undetermined reads and demultiplexing metrics should be reviewed before downstream analysis

## Public Tool Boundary

Prefer `bcl-convert` if it is already installed. It is free for local use but proprietary and RPM-distributed by Illumina, so do not auto-download without explicit user approval.

Legacy `bcl2fastq` may exist in older environments. Use it only when BCL Convert is unavailable or the run requires legacy compatibility.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bcl_to_fastq --emit-install-plan
```

Also check run-folder structure:

```bash
test -f /path/to/run/RunInfo.xml
test -f /path/to/SampleSheet.csv
find /path/to/run -maxdepth 4 -type d -name BaseCalls
```

## Local Execution Package

Use the plugin-owned runner when the user provides a local run folder and sample sheet:

```bash
python plugins/ngs-analysis/scripts/run_bcl_to_fastq.py \
  --run-folder /path/to/run \
  --sample-sheet /path/to/SampleSheet.csv \
  --output-directory /path/to/fastq_out
```

Add `--execute` only when conversion is requested. The runner validates `RunInfo.xml`, optional `RunParameters.xml`, the BaseCalls directory, sample-sheet rows, duplicate lane/index combinations, and index length compatibility. With `--execute`, it uses installed `bcl-convert`, then legacy `bcl2fastq` if available; if neither exists, it records the blocker instead of downloading proprietary software.

## Validation Checklist

Before conversion, validate:

- `RunInfo.xml` exists and its read structure matches the expected sequencing design.
- `SampleSheet.csv` exists, is the intended version, and has no duplicate sample/index combinations within each lane.
- Index sequence lengths match the index reads and any trimming/masking requested by the sample sheet.
- Dual-index orientation is explicit for the instrument and library prep; do not infer i5 orientation from filenames.
- UMI bases are assigned to the intended read or index read and carried through to FASTQ headers or output metadata as needed.
- Lane-splitting, sample-name normalization, and output directory behavior are agreed before running.
- Disk space is sufficient for output FASTQs, reports, and temporary files.

## Kickoff Pattern

First produce a preflight plan with paths and sample sheet validation. Then run conversion only after the user confirms:

```bash
bcl-convert \
  --bcl-input-directory /path/to/run \
  --output-directory /path/to/fastq_out \
  --sample-sheet /path/to/SampleSheet.csv
```

## Metrics Review

After conversion, inspect and report:

- total clusters, clusters passing filter, and yield by lane
- percent assigned by sample and percent undetermined by lane
- top undetermined index sequences when available
- per-sample FASTQ counts and read-pair consistency
- unexpected index hopping, barcode collision, or sample-sheet mismatch signals

Record software version, command, sample sheet checksum, run-folder path, output path, and conversion metrics. Do not start downstream analysis until severe demultiplexing anomalies are surfaced.

Referenced files: 1

ngs-bulk-rnaseq3.62 KB

View saved version →

---
name: ngs-bulk-rnaseq
description: Dispatch bulk RNA-seq requests to FASTQ-to-count QC or count-matrix differential-expression skills using nf-core/rnaseq, STAR, Salmon, featureCounts, MultiQC, and R/Bioconductor workflows.
---

# Bulk RNA-seq

Use this skill as the bulk RNA-seq dispatcher. Route FASTQ/BAM processing to count-generation QC, and route count-matrix statistical analysis to differential-expression guidance.

## Essential Inputs

Confirm:

- organism and genome build
- FASTA and GTF, or supported nf-core genome key
- paired-end or single-end reads
- strandedness, or whether to infer strandedness
- sample sheet and metadata
- counts-only vs differential expression
- contrasts, covariates, and batch terms for differential expression

## Dispatch

- FASTQ or aligned reads to raw counts, transcript estimates, or MultiQC summaries: `ngs-bulk-rnaseq-counts-qc`
- Raw count matrix plus sample metadata to contrasts, plots, and DE result tables: `ngs-bulk-rnaseq-differential-expression`

If the user asks for both, run count-generation planning first and start differential expression only after the raw count matrix, sample metadata, replicates, design formula, and contrasts are confirmed.

## Public Default

Prefer `nf-core/rnaseq` for standardized processing when a stable container or HPC runtime is available. Use the `local_light` Snakemake/Salmon path when Docker, registry egress, or Nextflow process containers are unavailable and a compact local run is appropriate.

## Plugin-Owned Local Paths

Use the counts/QC runner for local FASTQ-to-matrix execution:

```bash
python plugins/ngs-analysis/scripts/run_bulk_rnaseq_counts_qc.py \
  --sample-sheet samplesheet.csv \
  --fastq-root path/to/fastqs \
  --transcriptome-fasta reference/transcriptome.fasta \
  --genome-fasta reference/genome.fa \
  --annotation-gtf reference/genes.gtf \
  --execute
```

Use the differential-expression runner when the user already has a count or expression matrix:

```bash
python plugins/ngs-analysis/scripts/run_bulk_rnaseq_de.py \
  --count-matrix count_matrix.tsv \
  --sample-metadata sample_metadata.tsv \
  --contrasts contrasts.tsv \
  --execute
```

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_counts_qc --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_differential_expression --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --profile local_light --emit-install-plan
```

## Kickoff Pattern

Preflight run:

```bash
nextflow run nf-core/rnaseq \
  -profile test,docker \
  --outdir results/rnaseq_test
```

Real run skeleton:

```bash
nextflow run nf-core/rnaseq \
  -profile docker \
  --input samplesheet.csv \
  --outdir results/rnaseq \
  --genome GRCh38 \
  --aligner star_salmon
```

If strandedness is unknown, run inference or use the pipeline's strandedness detection before committing to final counts.

Local execution run:

```bash
python plugins/ngs-analysis/scripts/run_bulk_rnaseq_counts_qc.py \
  --sample-sheet samplesheet.csv \
  --fastq-root path/to/fastqs \
  --transcriptome-fasta reference/transcriptome.fasta
```

The local runners create a standard run envelope with `run_manifest.json`, `config.json`, `validation/`, `logs/`, `versions/`, `artifact_index.json`, and `summary.md`. Do not depend on development-only eval harness paths in a shared package.

## Downstream

Only start DESeq2/edgeR/limma analysis after confirming biological replicates, design formula, and contrasts. Preserve the raw count matrix and sample metadata.

Referenced files: 1

ngs-bulk-rnaseq-counts-qc3.94 KB

View saved version →

---
name: ngs-bulk-rnaseq-counts-qc
description: Run or plan bulk RNA-seq FASTQ-to-count processing with sample-sheet, strandedness, genome annotation, alignment or pseudoalignment, MultiQC, and count-matrix QC checks.
---

# Bulk RNA-seq Counts QC

Use this skill for bulk RNA-seq read processing, quantification, and count-matrix generation. If the user already has a count matrix and wants contrasts or statistics, use `ngs-bulk-rnaseq-differential-expression`.

## Essential Inputs

Confirm:

- FASTQ or aligned-read inputs and paired-end/single-end status
- organism, genome build, FASTA, GTF, and gene ID convention
- strandedness or permission to infer strandedness
- sample sheet with biological condition, replicate, batch, and library metadata
- desired quantification: gene counts, transcript estimates, or both
- alignment strategy: `STAR/Salmon`, Salmon-only, featureCounts from BAMs, or existing lab protocol

## Route

Prefer `nf-core/rnaseq` for standard processing when a stable container or HPC runtime is available. Use the `local_light` Snakemake/Salmon path for small local/devbox feasibility runs when Docker, registry egress, or Nextflow process containers are the blocker.

The plugin-owned local runner is:

```bash
python plugins/ngs-analysis/scripts/run_bulk_rnaseq_counts_qc.py \
  --sample-sheet samplesheet.csv \
  --fastq-root path/to/fastqs \
  --transcriptome-fasta reference/transcriptome.fasta \
  --genome-fasta reference/genome.fa \
  --annotation-gtf reference/genes.gtf \
  --execute
```

Omit `--execute` for validation plus Snakemake workflow validation only. Use `--no-dry-run` only when the user wants input validation and run-envelope preparation without workflow graph validation.

The runner emits a run-local `resources/` readiness bundle with `resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. Resource checks are advisory by default for custom or reduced references; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when a registered genome bundle must be complete before the run is considered ready.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_counts_qc --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --profile local_light --emit-install-plan
```

## Decision Points

- If strandedness is unknown, infer it before final counting; do not lock in a design based on library guesses.
- If strandedness is provided, carry it into the quantification command and flag any disagreement between the configured library type and Salmon's inferred format.
- Keep genome FASTA, GTF, transcriptome, and aligner indexes from the same build/release.
- Inspect per-sample reads, mapping rate, rRNA/mitochondrial fraction when available, duplication, insert size, gene-body bias, and assignment rate.
- Preserve raw counts separately from normalized expression.
- Carry sample metadata forward exactly; downstream DE depends on this table.

## Outputs

Produce:

- sample sheet and command/profile
- reference manifest with genome and GTF release
- MultiQC or equivalent processing summary
- Salmon `quant.sf` outputs, TPM/NumReads/effective-length matrices, and carried-forward sample metadata
- Gene-level expected-count and TPM matrices derived from transcript-level Salmon outputs, plus a `tx2gene` provenance table
- Compact QC verdict JSON covering mapping rate, duplication, library-type agreement, and outlier samples
- Browser-safe MultiQC helper HTML pages and a localhost launch hint for reliable in-app review
- Run-local reference readiness artifacts under `resources/`, including the resource plan, manifest, environment exports, and Markdown readiness summary
- issues that block differential expression, such as missing replicates, mislabeled groups, or severe batch/library failures
- standard run envelope: `run_manifest.json`, `config.json`, `validation/`, `logs/`, `versions/`, `artifact_index.json`, and `summary.md`

Referenced files: 1

ngs-bulk-rnaseq-differential-expression3.46 KB

View saved version →

---
name: ngs-bulk-rnaseq-differential-expression
description: Run or plan bulk RNA-seq differential-expression analysis from count matrices with replicate, design formula, contrast, batch, normalization, QC plot, and result-table checks.
---

# Bulk RNA-seq Differential Expression

Use this skill when the user has raw counts or a count-generation output and wants differential expression, contrasts, QC plots, or ranked gene tables.

## Essential Inputs

Confirm:

- raw count matrix path and sample metadata path
- gene ID type and annotation mapping requirement
- biological conditions, replicates, batch variables, donor pairing, covariates, and exclusions
- exact contrasts and baseline levels
- preferred statistical framework: DESeq2, edgeR, limma-voom, or existing lab standard
- output needs: normalized counts, PCA, sample distance, volcano plots, heatmaps, ranked tables, GSEA-ready lists

## Preconditions

Do not start differential expression until:

- raw counts are preserved
- each requested contrast has enough biological replication
- sample metadata row names match count matrix columns
- batch/covariate choices are explicit
- exploratory PCA/sample-distance plots do not reveal obvious swaps or failed libraries

## Route

For most count matrices, use DESeq2 or edgeR. Use limma-voom when the study design or lab standard favors it. Keep the analysis in R when using Bioconductor unless the user specifically asks for a Python-only workflow.

The plugin-owned local runner is:

```bash
python plugins/ngs-analysis/scripts/run_bulk_rnaseq_de.py \
  --count-matrix count_matrix.tsv \
  --sample-metadata sample_metadata.tsv \
  --contrasts contrasts.tsv \
  --execute
```

Use `--method auto` unless the user or lab standard specifies `DESeq2`, `edgeR`, or `limma_log2`. Auto mode uses DESeq2 when integer-like counts and the package are available, falls back to edgeR for integer-like counts, and uses `limma_log2` for non-integer expression matrices.

Use `--input-mode` to declare whether the matrix is `raw_counts`, `normalized_expression`, or `log_expression`. When `--input-mode auto` is used, the runner infers the mode and records a warning if normalization is skipped because the matrix is already transformed.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_differential_expression --emit-install-plan
```

## Decision Points

- Never compare groups without stating the design formula and contrast.
- Treat batch correction in modeling separately from visual batch removal.
- Do not filter genes using post-hoc knowledge of the contrast.
- For paired or repeated-measures designs, model subject/donor explicitly.
- Report genes with effect size, uncertainty, adjusted p-value, and filtering status.

## Outputs

Produce:

- design formula and contrast manifest
- QC plots: library size, detected genes, PCA/sample distance, mean-variance trend, and outlier review
- input-mode-aware matrix exports plus the modeling/log-scale matrix used for DE
- differential-expression tables per contrast
- explicit `.not_tested.tsv` stubs for contrasts blocked by insufficient replication or confounding
- auto-launched localhost Marimo review app recorded in `notebooks/marimo_server.json`
- caveats for small n, confounded designs, failed samples, or batch variables that cannot be estimated
- standard run envelope: `run_manifest.json`, `config.json`, `validation/`, `logs/`, `versions/`, `visualizations/`, `notebooks/`, `artifact_index.json`, and `summary.md`

Referenced files: 1

ngs-chip-cutrun-peaks-qc3.67 KB

View saved version →

---
name: ngs-chip-cutrun-peaks-qc
description: Run or plan ChIP-seq, CUT&RUN, or CUT&Tag QC, control handling, spike-in, peak calling, broad-vs-narrow target selection, replicate, bigWig, and differential binding workflows.
---

# ChIP/CUT&RUN Peaks QC

Use this skill for antibody-targeted enrichment workflows: ChIP-seq, CUT&RUN, or CUT&Tag. Use `ngs-atacseq-peaks-qc` for ATAC-seq.

## Essential Inputs

Confirm:

- assay: ChIP-seq, CUT&RUN, or CUT&Tag
- target class: transcription factor, histone mark, chromatin regulator, or custom target
- FASTQ/BAM inputs and paired-end status
- input DNA, IgG, no-antibody, or spike-in controls
- organism, genome build, blacklist, and spike-in genome if used
- biological replicates, conditions, batches, and sample metadata
- desired endpoint: QC, peaks, bigWigs, consensus peaks, or differential binding

## Route

Use `nf-core/chipseq` for ChIP-seq and `nf-core/cutandrun` for CUT&RUN/CUT&Tag when they fit the assay. Use direct MACS2 only for prepared BAMs with known control and duplicate policy.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline chip_cutrun_peaks_qc --emit-install-plan
```

For compact FASTQ intake/QC, use the shared epigenomics execution package:

```bash
python plugins/ngs-analysis/scripts/run_fastq_assay_package.py \
  --lane epigenomics_peaks \
  --sample-sheet chip_or_cutrun_samples.csv \
  --execute
```

It records FASTQ-level QC and peak-calling readiness.

For local-light alignment, control-aware MACS2 peak calling, FRiP, bigWig tracks, consensus peaks, and motif-handoff artifacts, use the dedicated ChIP/CUT&RUN runner:

```bash
python plugins/ngs-analysis/scripts/run_chip_cutrun_peaks_qc.py \
  --sample-sheet chip_or_cutrun_samples.csv \
  --assay chipseq \
  --target-class tf \
  --peak-mode narrow \
  --bowtie2-index /refs/GRCh38/bowtie2/genome \
  --genome-size hs \
  --blacklist-bed /refs/GRCh38/blacklists/encode_blacklist.bed \
  --execute
```

This runner emits `qc/chip_cutrun_qc_summary.{tsv,json}`, `qc/chip_cutrun_qc_dashboard.html`, native SVG FRiP/peak and insert-size plots, browser-track handoff files under `tracks/`, and `motifs/motif_summary.tsv`. Add `--run-motifs --motif-genome <genome>` when HOMER motif enrichment should be executed instead of only planned.

It also emits `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. The resource check is advisory by default for local-light runs; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when missing registered reference bundles should block readiness.

For nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline chipseq` or `--pipeline cutandrun`.

## Decision Points

- Choose narrow versus broad peak mode from target biology, not from convenience.
- Preserve control pairing and spike-in metadata through sample sheets.
- For histone marks, expect broad or domain-like signal for many marks; for TFs, expect sharper peaks and stronger replicate checks.
- Review alignment rate, duplicate rate, fragment size, FRiP/peak signal, blacklist overlap, and replicate concordance.
- Keep consensus peak generation and differential binding design separate from raw peak calling.

## Outputs

Produce:

- assay/target/control manifest
- command/profile and sample sheet
- QC summary with replicate/control status
- peaks, bigWigs, browser-track manifests, browser-track preview HTML, native QC dashboard/SVG plots, consensus peaks, and count matrix when requested
- motif summary files when a motif backend is requested
- differential binding design and caveats for missing controls, weak enrichment, or poor replicate concordance

Referenced files: 1

ngs-dna-germline-variants3.72 KB

View saved version →

---
name: ngs-dna-germline-variants
description: Run or plan deep germline WGS, WES, targeted-panel, cohort, or trio variant-calling workflows with reference-build, known-sites, QC, joint-calling, and annotation checks.
---

# Germline DNA Variants

Use this skill for germline WGS, WES, or inherited-disease panel analysis from FASTQ, BAM, or CRAM. If the request is tumor-only, tumor-normal, or low-frequency molecular-barcode panel calling, use a somatic or UMI-panel skill instead.

## Essential Inputs

Confirm:

- data type: WGS, WES, or targeted panel
- sample model: singleton, cohort, duo, trio, family, or case/control
- input type: FASTQ, BAM, or CRAM
- organism, reference build, FASTA, indexes, and contig naming
- known-sites resources for BQSR, contamination, and annotation
- target BED and bait BED for WES/panel data
- sex/ploidy assumptions and mitochondrial/sex-chromosome requirements
- desired callers, annotation outputs, and final VCF/gVCF expectations

## Route

Prefer `nf-core/sarek` for full FASTQ/BAM-to-VCF workflows. Use direct GATK4, DeepVariant, samtools, or bcftools only for focused tasks or a custom workflow.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_germline_variants --emit-install-plan
```

For compact local checks from prepared BAM/CRAM files, use the shared DNA execution package:

```bash
python plugins/ngs-analysis/scripts/run_dna_variant_calling.py \
  --sample-sheet dna_samples.tsv \
  --reference-fasta reference.fa \
  --execute
```

Treat this as a focused samtools/bcftools run envelope, not as a substitute for full cohort, trio, gVCF, BQSR, or annotation workflows.

For a higher-fidelity local germline run that owns BQSR, per-sample gVCFs, and joint genotyping assumptions, use the germline-specific runner:

```bash
python plugins/ngs-analysis/scripts/run_dna_germline_variants.py \
  --sample-sheet dna_samples.tsv \
  --reference-fasta reference.fa \
  --known-sites dbsnp.vcf.gz \
  --known-sites mills.vcf.gz \
  --emit-gvcf \
  --joint-call \
  --execute
```

This runner still expects reference-matched resources and an available GATK toolchain. It packages the validation state and generated artifacts even when execution is blocked by missing tools or resources.

It also writes advisory `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md` artifacts by default. Add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when complete registered reference and known-sites bundles should be mandatory for readiness.

## Decision Points

- For cohorts or families, decide whether the endpoint is per-sample VCFs, gVCFs for joint genotyping, or a jointly called cohort VCF.
- For WES/panels, carry the target BED through alignment metrics, calling, and coverage reports; do not call off-target regions by accident.
- Use BQSR only when reference-matched known-sites resources exist. Do not mix GRCh37, hg19, GRCh38, or T2T resources.
- Check sample identity, sex concordance, contamination, coverage, duplication, insert size, and transition/transversion where feasible.
- For trios, preserve pedigree metadata and report Mendelian/QC checks separately from variant interpretation.

## Outputs

Produce:

- command or workflow profile and sample sheet
- reference/resource manifest with versions and checksums when available
- QC summary: coverage, duplication, insert size, contamination, sex/relatedness checks when run
- VCF/gVCF path, index path, and annotation path
- limitations: low coverage, missing known-sites, target design gaps, or build mismatches

Clinical interpretation, pathogenicity classification, and report signing are out of scope unless the user provides a validated clinical workflow.

Referenced files: 1

ngs-dna-somatic-variants3.72 KB

View saved version →

---
name: ngs-dna-somatic-variants
description: Run or plan tumor-normal, tumor-only, WGS, WES, or cancer-panel somatic variant workflows with pairing, contamination, panel-of-normals, purity, QC, and annotation checks.
---

# Somatic DNA Variants

Use this skill for tumor-normal or tumor-only somatic SNV/indel calling from FASTQ, BAM, or CRAM. If the request is inherited germline calling or family analysis, use `ngs-dna-germline-variants`.

## Essential Inputs

Confirm:

- tumor-normal, tumor-only, relapse-baseline, or multi-tumor design
- WGS, WES, or panel assay and target BED when applicable
- input type and whether reads are already aligned
- tumor/normal pairing table and sample identifiers
- reference build, known-sites, germline resource, and annotation cache
- panel-of-normals availability and matched-normal availability
- tumor purity, contamination expectations, and minimum allele fraction goals
- desired outputs: raw calls, filtered calls, VEP/SnpEff annotation, MAF, CNV/SV handoff

## Route

Prefer `nf-core/sarek` for an end-to-end public workflow when its supported callers fit the request. Use direct GATK Mutect2 or bcftools/samtools utilities for focused validation or prepared BAMs.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_somatic_variants --emit-install-plan
```

For compact local checks from prepared tumor/normal BAM/CRAM files, use the dedicated Mutect2 runner:

```bash
python plugins/ngs-analysis/scripts/run_dna_somatic_variants.py \
  --sample-sheet somatic_pairs.tsv \
  --reference-fasta reference.fa \
  --germline-resource af-only-gnomad.vcf.gz \
  --panel-of-normals pon.vcf.gz \
  --execute
```

This produces a tumor-normal/tumor-only pairing table, Mutect2 command plan, contamination/filtering artifacts, somatic QC summary, `qc/somatic_pair_review.{tsv,json}`, visualization index, and filtered VCF outputs when the local GATK resources are available. For nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline sarek`.

The direct runner also emits `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. The resource check is advisory by default so custom or reduced references can still be planned; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when missing registered reference bundles should block readiness.

## Decision Points

- Verify tumor-normal pair metadata before execution. A swapped or missing normal changes the biological meaning of the calls.
- For tumor-only analysis, explicitly state the false-positive risk and require a germline resource plus careful filtering.
- Use panel-of-normals when available and reference-matched; do not reuse a PON across incompatible capture kits or genome builds.
- Track contamination, orientation bias, strand artifacts, mapping quality, coverage, tumor purity, and allele-fraction filters.
- Keep germline filtering separate from somatic interpretation; avoid presenting tumor-only calls as confirmed somatic without supporting evidence.

## Outputs

Produce:

- validated pairing/sample sheet
- caller/filter settings and reference/resource manifest
- QC summary: tumor/normal depth, contamination, duplication, insert size, on-target rate for panels/WES
- per-pair review table covering matched-normal state, PON/germline-resource availability, contamination-table status, filtered VCF status, and parsed variant counts
- VCF/MAF/annotation paths and a filtered-vs-raw call count summary
- caveats for tumor-only calls, low-purity tumors, low-depth regions, or missing matched normals

Clinical actionability and treatment recommendations are out of scope unless the user supplies a validated clinical interpretation workflow.

Referenced files: 1

ngs-dna-umi-panel-variants3.71 KB

View saved version →

---
name: ngs-dna-umi-panel-variants
description: Run or plan targeted DNA panel variant workflows that use UMIs, duplex consensus reads, molecular barcodes, low-frequency calling, target coverage, and panel-specific QC.
---

# UMI Panel DNA Variants

Use this skill for targeted DNA panels where molecular barcodes, UMIs, duplex consensus, or low-frequency allele detection are central to the analysis. If the panel is ordinary germline calling without molecular consensus, use `ngs-dna-germline-variants`.

## Essential Inputs

Confirm:

- panel/capture kit name and target BED
- UMI layout: inline read, index read, single UMI, duplex UMI, or unknown
- whether consensus reads have already been generated
- FASTQ/BAM input and pairing convention
- reference build and panel-specific annotation requirements
- minimum allele fraction goal and intended use: screening, research, validation, or exploratory
- positive/negative controls and expected spike-ins when available

## Route

Use a lab-validated panel workflow when provided. For public-tool planning, combine FASTQ QC, UMI extraction/consensus generation, alignment, target coverage QC, and variant calling as separate audited stages.

Preflight command:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_umi_panel_variants --emit-install-plan
```

For compact local checks from prepared consensus or alignment BAM/CRAM files, use the dedicated UMI panel runner:

```bash
python plugins/ngs-analysis/scripts/run_dna_umi_panel_variants.py \
  --sample-sheet umi_panel_samples.tsv \
  --reference-fasta reference.fa \
  --target-bed panel_targets.bed \
  --umi-mode duplex \
  --umi-tag RX \
  --execute
```

This writes the consensus/variant command plan, molecular-consensus state, low-frequency calling settings, visualization index, `qc/umi_postrun_summary.{tsv,json}`, `qc/umi_molecular_evidence_contract.{tsv,json}`, and consensus-BAM VCF outputs when the local fgbio/samtools/bcftools backend is available. The post-run summary parses consensus flagstat, target coverage, bcftools stats, and family-size/duplex files when present; missing metrics stay explicit in the notes column. The molecular evidence contract keeps the low-AF review requirements visible per sample: consensus BAM, family-size or molecule-support metrics, variant stats, hotspot review, and duplex review.

The direct runner also emits `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. The resource check is advisory by default so custom or reduced references can still be planned; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when missing registered reference bundles should block readiness.

## Decision Points

- Do not trim or discard UMI bases until their layout and destination are known.
- Separate raw read depth from unique molecular depth and consensus depth.
- Track on-target rate, coverage uniformity, family size distribution, strand/duplex support, and per-target dropout.
- Low allele fraction calls require stronger artifact review than ordinary germline calls.
- Use panel-specific hotspot/blacklist rules only when their provenance is known.

## Outputs

Produce:

- UMI layout and consensus strategy
- target BED/resource manifest
- raw-depth, molecular-depth, and consensus-depth QC summary
- `qc/umi_postrun_summary.tsv` for consensus reads, target coverage, variant counts, family size, and duplex fraction
- `qc/umi_molecular_evidence_contract.tsv` for low-AF evidence readiness, hotspot review, and duplex review expectations
- variant calls with allele fraction, depth, strand/duplex support, and filtering rationale
- limitations around sensitivity, panel dropout, molecule count, and non-validated interpretation

Referenced files: 1

ngs-dna-variant-calling3.67 KB

View saved version →

---
name: ngs-dna-variant-calling
description: Dispatch WGS, WES, or targeted DNA variant requests to germline, somatic, or UMI-panel skills, then plan public nf-core/sarek, GATK4, DeepVariant, samtools, or bcftools workflows.
---

# DNA Variant Calling

Use this skill as the DNA variant-calling dispatcher for WGS, WES, or targeted DNA panel analysis from FASTQ, BAM, or CRAM. Once the sample model is clear, hand off to the narrow subtype skill.

## Essential Inputs

Confirm:

- data type: WGS, WES, or panel
- sample model: germline single sample, cohort, trio, tumor-only, or tumor-normal
- input type: FASTQ, BAM, or CRAM
- organism and reference genome
- known-sites resources for BQSR, if required
- target BED for WES or panels
- UMI or duplex handling
- desired callers and annotation outputs

## Dispatch

Route by biological/sample model:

- Germline singleton, cohort, family, trio, WGS, WES, or ordinary inherited panel: `ngs-dna-germline-variants`
- Tumor-normal, tumor-only, relapse-baseline, or other cancer somatic calling: `ngs-dna-somatic-variants`
- UMI, duplex, molecular-barcode, or low-frequency targeted panel calling: `ngs-dna-umi-panel-variants`

If the request is ambiguous, ask only for the missing sample model and assay design needed to choose among these three. Do not run one generic variant workflow when the request needs subtype-specific assumptions.

## Public Default

Prefer `nf-core/sarek` for an end-to-end public workflow. Use direct GATK4, DeepVariant, samtools, or bcftools commands only for smaller, focused tasks or when the user explicitly wants a custom pipeline.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_variant_calling --emit-install-plan
```

## Local Execution Package

For a compact BAM/CRAM-to-VCF run with a matching reference FASTA, use the plugin-owned samtools/bcftools runner:

```bash
python plugins/ngs-analysis/scripts/run_dna_variant_calling.py \
  --sample-sheet dna_samples.tsv \
  --reference-fasta reference.fa \
  --region chr20:1-100000 \
  --filter-min-qual 30 \
  --filter-min-site-dp 10 \
  --execute
```

The sample sheet should include `sample` and `bam` or `cram` columns. When `--region` is provided the runner also emits per-base depth plus a callable-loci summary for that interval, and when filter thresholds are provided it emits a soft-filtered VCF alongside the raw calls. This package is suitable for focused local checks and run-envelope generation; subtype skills still own germline, somatic, UMI, reference-resource, cohort, annotation, and workflow assumptions.

This compact runner now writes advisory `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md` artifacts for the selected genome bundle. Use `--require-resource-plan` when missing registered reference resources should block readiness; otherwise the explicit `--reference-fasta` remains enough for focused local checks.

## Kickoff Pattern

Preflight-first nf-core pattern:

```bash
nextflow run nf-core/sarek \
  -profile test,docker \
  --outdir results/sarek_test
```

Real run skeleton:

```bash
nextflow run nf-core/sarek \
  -profile docker \
  --input samplesheet.csv \
  --outdir results/sarek \
  --genome GRCh38 \
  --tools haplotypecaller,vep
```

For WES/panel data, include the target BED. For tumor-normal data, verify pair metadata before execution. For UMI panels, preserve barcode handling and molecule-level QC.

## Guardrails

- Do not mix genome builds across FASTA, GTF/BED, known sites, and VEP cache.
- Do not download large references without confirming disk space and target path.
- Treat clinical interpretation as out of scope unless the user has a validated clinical workflow.

Referenced files: 1

ngs-epigenomics-peaks2.5 KB

View saved version →

---
name: ngs-epigenomics-peaks
description: Dispatch ATAC-seq, ChIP-seq, CUT&RUN, or CUT&Tag requests to assay-specific QC, alignment, signal-track, peak-calling, consensus, and differential peak workflows.
---

# Epigenomics Peaks

Use this skill as the epigenomics dispatcher for ATAC-seq, ChIP-seq, CUT&RUN, or CUT&Tag analysis. Hand off to the assay-specific deep skill once the assay type is known.

## Essential Inputs

Confirm:

- assay type
- FASTQ or BAM input
- organism and genome build
- blacklist file, if available
- control samples: input DNA, IgG, or spike-in
- biological replicates
- peak type: narrow, broad, accessibility, or protocol-specific
- desired outputs: QC report, peaks, consensus peaks, bigWigs, differential peaks

## Public Defaults

Choose the workflow by assay:

- ATAC-seq: `ngs-atacseq-peaks-qc` using `nf-core/atacseq` by default
- ChIP-seq: `ngs-chip-cutrun-peaks-qc` using `nf-core/chipseq` by default
- CUT&RUN or CUT&Tag: `ngs-chip-cutrun-peaks-qc` using `nf-core/cutandrun` by default

Use direct MACS2 only for focused peak-calling tasks from prepared BAMs.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline epigenomics_peaks --emit-install-plan
```

## Local Execution Package

For FASTQ intake/QC over ATAC-seq, ChIP-seq, CUT&RUN, or CUT&Tag data, use the shared FASTQ assay package:

```bash
python plugins/ngs-analysis/scripts/run_fastq_assay_package.py \
  --lane epigenomics_peaks \
  --sample-sheet assay_samples.csv \
  --execute
```

This validates sample-sheet paths and read structure, runs seqkit stats and FastQC/MultiQC when available, and writes `peak_calling_readiness.json`. Full alignment, signal tracks, TSS/FRiP, consensus peaks, and differential analyses still route through the assay-specific workflow.

Assay-specific ATAC and ChIP/CUT&RUN runners now also emit native review files alongside TSV/JSON summaries: `qc/*_dashboard.html`, FRiP/peak SVG plots, insert-size SVG plots, browser-track preview HTML, UCSC track lines, and IGV session files.

## Kickoff Pattern

ATAC-seq preflight run:

```bash
nextflow run nf-core/atacseq \
  -profile test,docker \
  --outdir results/atacseq_test
```

ChIP-seq preflight run:

```bash
nextflow run nf-core/chipseq \
  -profile test,docker \
  --outdir results/chipseq_test
```

CUT&RUN/CUT&Tag preflight run:

```bash
nextflow run nf-core/cutandrun \
  -profile test,docker \
  --outdir results/cutandrun_test
```

Carry replicate and control metadata through the sample sheet before running real analysis.

Referenced files: 1

ngs-fastq-qc4.28 KB

View saved version →

---
name: ngs-fastq-qc
description: Validate FASTQ inputs, run local FastQC/MultiQC QC, interpret QC signals, and optionally execute fastp or Cutadapt trimming branches without overwriting raw reads.
---

# FASTQ QC

Use this skill for QC-only, trimming-first, or FASTQ quality interpretation workflows. This skill can execute the plugin-owned local FastQ QC runner when the user approves a local run. It should decide whether trimming or additional investigation is warranted; it should not blindly trim by default.

## Essential Inputs

Confirm:

- FASTQ paths and pairing convention
- whether output should be QC-only or trimmed FASTQs
- known adapter or primer sequences
- organism if contamination screening or host depletion is requested
- output directory
- whether FASTQs are raw, demultiplexed, previously trimmed, or downloaded from an archive
- whether downstream analysis expects original read lengths, UMIs, or inline barcodes

## Public Tools

Default tool set:

- `FastQC` for raw read QC
- `MultiQC` for project-level summary
- `fastp` for all-in-one QC/trimming when acceptable
- `Cutadapt` when primer/adapter handling needs explicit sequences
- `seqkit` for quick counts, stats, and subsampling

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline fastq_qc --emit-install-plan
```

## Local Execution

Use the plugin-owned runner for local artifact-producing FASTQ QC:

```bash
python plugins/ngs-analysis/scripts/run_fastq_qc.py \
  --sample-sheet samplesheet.csv \
  --execute
```

Single paired sample:

```bash
python plugins/ngs-analysis/scripts/run_fastq_qc.py \
  --sample sampleA \
  --r1 sampleA_R1.fastq.gz \
  --r2 sampleA_R2.fastq.gz \
  --execute
```

Optional trimming branch:

```bash
python plugins/ngs-analysis/scripts/run_fastq_qc.py \
  --sample-sheet samplesheet.csv \
  --trim-mode fastp \
  --execute
```

For explicit adapters:

```bash
python plugins/ngs-analysis/scripts/run_fastq_qc.py \
  --sample-sheet samplesheet.csv \
  --trim-mode cutadapt \
  --adapter-r1 AGATCGGAAGAGC \
  --adapter-r2 AGATCGGAAGAGC \
  --execute
```

The runner performs pre-execution validation before Snakemake execution. It writes a timestamped run directory with `run_manifest.json`, `config.json`, `validation/`, `workflow/Snakefile`, logs, `artifact_index.json`, `summary.md`, FastQC/MultiQC outputs, and `qc_interpretation.json` after successful execution.

## Interpretation Rules

Inspect raw QC before recommending trimming:

- Per-base quality drop at the read end: consider quality trimming, but preserve enough length for alignment or amplicon merging.
- Adapter or primer signal: use `cutadapt` when explicit sequences matter; use `fastp` only when automatic handling is acceptable.
- Poly-G or patterned-flowcell artifacts: handle with a tool that explicitly supports the artifact and report the assumption.
- Overrepresented sequences: classify adapters, primers, rRNA, PhiX, host contamination, or true biology before filtering.
- Per-tile failures or severe quality shifts: flag possible run-level issues and avoid treating them as ordinary adapter contamination.
- High duplication: interpret by assay; it may be expected for amplicons, targeted panels, or low-input libraries.
- Pairing issues: verify R1/R2 file counts and read-name pairing before any downstream workflow.

Do not overwrite input FASTQs. Preserve the raw QC reports even when trimmed FASTQs are created.

## Kickoff Pattern

QC-only:

```bash
mkdir -p results/fastqc results/multiqc
fastqc -t 4 -o results/fastqc *.fastq.gz
multiqc results/fastqc -o results/multiqc
```

QC plus trimming:

```bash
fastp \
  -i sample_R1.fastq.gz \
  -I sample_R2.fastq.gz \
  -o results/trimmed/sample_R1.fastq.gz \
  -O results/trimmed/sample_R2.fastq.gz \
  --html results/fastp/sample.html \
  --json results/fastp/sample.json
multiqc results -o results/multiqc
```

## Output Review

Return a short QC interpretation with:

1. sample/read-pair inventory
2. QC modules that look normal
3. QC modules that require action or user confirmation
4. trimming or no-trimming recommendation with rationale
5. downstream caveats such as short reads, contaminated libraries, or failed pairs

When using the local runner, ground the response in the generated `qc_interpretation.json`, `summary.md`, and MultiQC report instead of relying only on expected artifacts.

Referenced files: 1

ngs-runtime-env6.23 KB

View saved version →

---
name: ngs-runtime-env
description: Check whether public NGS tools and packages already exist before downloading, installing, or running a sequencing pipeline.
---

# NGS Runtime Environment

Use this skill whenever an NGS workflow needs package checks, install planning, or runtime validation.

## Existence Check Order

1. Check executables on `PATH` with `command -v` or `shutil.which`.
2. Check Python imports for Python-backed tools.
3. Check active package managers with `conda list`, `mamba list`, `micromamba list`, or `pip show`.
4. If requested, check package indexes or container registries.
5. Emit an install plan before installing.
6. Install only when explicitly requested by the user.

Do not modify system Python. Prefer isolated conda/mamba environments or containers.

## Script

From the repo root:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --list
python plugins/ngs-analysis/scripts/ngs_preflight.py --tool fastqc --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --profile local_light --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_variant_calling --network-checks --emit-install-plan
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline shotgun_metagenomics --manager micromamba --install-plan-outdir runtime_readiness/shotgun_install
```

Use `--install-plan-outdir` when a user needs a reviewable permission handoff. It writes `install_plan.json` as the canonical machine-readable plan and `install_commands.sh` as a guarded shell companion generated from the same plan. The shell companion is review-only by default; it exits without installing unless `NGS_RUN_INSTALL_COMMANDS=1` is set after explicit user approval.

Check reference and database bundle readiness separately from executable readiness:

```bash
python plugins/ngs-analysis/scripts/ngs_reference_manager.py list
python plugins/ngs-analysis/scripts/ngs_reference_manager.py check --kind reference --bundle grch38_core --root /refs/GRCh38
python plugins/ngs-analysis/scripts/ngs_reference_manager.py explain-missing --kind database --bundle kraken2_standard --root /db/kraken2/standard
python plugins/ngs-analysis/scripts/ngs_reference_manager.py plan --pipeline shotgun_metagenomics --include-optional --outdir resource_readiness/shotgun
python plugins/ngs-analysis/scripts/ngs_reference_manager.py setup-plan --pipeline shotgun_metagenomics --include-optional --outdir resource_readiness/shotgun_setup
python plugins/ngs-analysis/scripts/ngs_reference_manager.py plan --pipeline atacseq --genome-build GRCh38 --bundle-root grch38_core=/refs/GRCh38 --outdir resource_readiness/atac
python plugins/ngs-analysis/scripts/ngs_reference_manager.py inventory --outdir resource_readiness/inventory
python plugins/ngs-analysis/scripts/ngs_reference_manager.py lock --outdir resource_readiness/lock --include-checksums
python plugins/ngs-analysis/scripts/ngs_reference_manager.py verify-lock --lockfile resource_readiness/lock/resource_lock.json --outdir resource_readiness/lock_verify --fail-on-mismatch
python plugins/ngs-analysis/scripts/ngs_reference_manager.py check-all --kind database --output resource_readiness/database_audit.json
```

Use `plan` before claiming that a reference- or database-heavy workflow is runnable. The plan output writes `resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, `resource_readiness.md`, and setup-plan artifacts; missing required bundles are blocking, while optional bundles such as Bracken/HUMAnN or HOMER motif resources should stay explicit.

Use `setup-plan` when the user needs an actionable resource/database setup checklist without running an assay. It writes `resource_setup_plan.json`, `resource_setup_plan.tsv`, `resource_setup_plan.md`, and `resource_setup_commands.sh`. The shell skeleton keeps setup hints commented by default, so large reference/database downloads remain deliberate and reviewable.

Use `inventory` when the user needs a broader resource/database audit across the plugin. It writes `resource_inventory.json`, `resource_inventory.tsv`, `resource_env.sh`, and `resource_dashboard.md`, including missing files, env vars, setup hints, license notes, and pipeline usage for every known bundle.

Use `lock` after resources are ready for a project or handoff. It snapshots the resource inventory into `resource_lock.json`, `resource_lock.tsv`, and `resource_lock.md`; `verify-lock` compares the lockfile against current local paths and writes a drift report before reruns.

The nf-core adapter performs the same resource gate automatically unless `--skip-resource-plan` is supplied:

```bash
python plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline taxprofiler --sample-sheet samples.csv --profile docker --bundle-root kraken2_standard=/db/kraken2/standard --include-optional-resources
```

The direct bulk RNA-seq counts/QC, scRNA FASTQ-to-count, generic DNA, germline DNA, somatic DNA, UMI panel, ATAC, ChIP/CUT&RUN, amplicon, and shotgun backend runners also emit run-local `resources/` readiness bundles. These direct runners use advisory resource checks by default so custom or reduced local inputs can still be planned; add `--require-resource-plan` when missing registered bundles should block readiness.

Use `--install-missing --yes` only after explicit user approval:

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline fastq_qc --manager mamba --install-missing --yes
```

## Install Strategy

Prefer these patterns:

- nf-core workflows: install/check `nextflow`; use Docker/Singularity/Apptainer profiles for process tools.
- local execution: install/check `snakemake`; use `mamba` or `micromamba` environments and avoid containers by default.
- small QC tools: install with `mamba` or `micromamba` from `conda-forge` and `bioconda`.
- Python analysis packages: install in a dedicated environment, not global Python.
- large databases and references: estimate size and check existing paths before downloading.
- pipeline resource plans: use `--bundle-root bundle=/path` or the registry `root_env` variables so downstream runs can cite the exact local bundle roots.

## Report

Summarize:

- present tools and paths
- missing tools
- package-index checks, if performed
- suggested install commands
- tools that are proprietary, EULA-bound, cloud-bound, or database-heavy

Referenced files: 1

ngs-scrna-seq3.78 KB

View saved version →

---
name: ngs-scrna-seq
description: Route single-cell or single-nucleus RNA-seq FASTQs to public count-generation workflows and defer post-count matrix QC, annotation, clustering, and UMAP analysis to the embedded scrna-seq-qc skill.
---

# Single-cell RNA-seq

Use this skill for scRNA-seq or snRNA-seq kickoff from FASTQs, Cell Ranger-style outputs, matrices, `.h5`, `.h5ad`, or `.rds`. This skill owns upstream intake and FASTQ-to-count routing; post-count QC, annotation, clustering, and UMAPs must route to the embedded `scrna-seq-qc` skill.

## Essential Inputs

Confirm:

- input type: FASTQ, count matrix, `.h5`, `.h5ad`, or `.rds`
- assay: single-cell or single-nucleus
- chemistry or barcode/UMI layout
- organism and reference
- expected cells per sample when available
- sample, donor, batch, and channel metadata
- desired endpoint: count matrix only, QC, clustering, annotation, UMAP, or differential abundance/expression

## Public Default

For FASTQs, prefer public alternatives:

- `nf-core/scrnaseq`
- STARsolo
- kallisto-bustools via `kb-python`
- alevin-fry

Use 10x Cell Ranger only when the user explicitly wants vendor-standard output and has accepted the 10x EULA.

## Implementation Sequence

Treat scRNA as three ordered rows in the plugin state and execute them sequentially:

1. FASTQ-to-count:
   count matrix generation, barcode and feature tables, chemistry or whitelist choice, and a backend summary.
2. Post-count QC and annotation:
   raw-count-preserving objects, QC metrics, threshold plots, doublet and ambient-RNA outputs, clustering, UMAPs, and annotation confidence.
3. Downstream stats:
   pseudobulk matrices, differential expression or abundance tables, and per-condition plots.

Cell Ranger is an optional backend when vendor-standard output is explicitly required. It is not a standalone roadmap row and it is not the default execution target.

For post-count QC/annotation, use the embedded `skills/scrna-seq-qc` guidance. Route to that skill whenever the requested endpoint starts from a matrix, `.h5`, `.h5ad`, `.rds`, Cell Ranger output, or asks for QC, doublets, ambient RNA, annotation, clustering, UMAPs, or post-count differential summaries.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline scrnaseq --emit-install-plan
```

## Kickoff Pattern

nf-core preflight run:

```bash
python plugins/ngs-analysis/scripts/run_nfcore_pipeline.py \
  --pipeline scrnaseq \
  --sample-sheet samplesheet.csv \
  --profile docker \
  --genome GRCh38 \
  --bundle-root grch38_core=/refs/GRCh38
```

This adapter captures the generated params, pinned Nextflow command, resource gate, trace/report paths, run manifest, and visualization index in the standard plugin envelope. Add `--revision <tag>` for pinned nf-core execution and `--execute` only when Nextflow plus a container/HPC profile are ready.

Plugin-owned local execution:

```bash
python plugins/ngs-analysis/scripts/run_scrnaseq_fastq_to_count.py \
  --sample-sheet samplesheet.csv \
  --genome-fasta reference/genome.fa \
  --annotation-gtf reference/genes.gtf \
  --cb-whitelist reference/whitelist.txt \
  --execute
```

The FASTQ-to-count runner emits advisory `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md` outputs by default. Add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when STARsolo reference bundle completeness should block readiness.

Matrix-level QC should be handled by `scrna-seq-qc` and must preserve raw counts, per-sample metadata, filter decisions, doublet calls, ambient-RNA handling, and plot outputs.

## Guardrails

- Do not assume 10x chemistry from filenames alone.
- Do not silently skip doublet or ambient-RNA assessment when doing QC.
- Do not over-annotate clusters without matched references or clear markers.

Referenced files: 1

ngs-shotgun-metagenomics4.77 KB

View saved version →

---
name: ngs-shotgun-metagenomics
description: Kick off public shotgun metagenomics QC, host-depletion, taxonomic profiling, and functional profiling workflows using nf-core/taxprofiler, Kraken2, Bracken, MetaPhlAn, and HUMAnN.
---

# Shotgun Metagenomics

Use this skill for shotgun metagenomic FASTQs.

## Essential Inputs

Confirm:

- paired-end or single-end reads
- host organism and host-depletion requirement
- target outputs: taxonomic profile, functional profile, assembly, binning, or QC only
- preferred database family, if any
- database paths or permission to download large databases
- sample metadata, batches, and negative controls

## Public Defaults

Prefer `nf-core/taxprofiler` for reproducible taxonomic profiling. Use direct Kraken2/Bracken, MetaPhlAn, or HUMAnN when the user wants a focused path or already has databases installed.

For direct backend execution, prefer the plugin runner over handwritten shell when possible because it validates database bundle contents and records `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. `--run-bracken` and `--run-humann` make those database bundles blocking, not merely optional.

## Preflight

```bash
python plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline shotgun_metagenomics --emit-install-plan
```

## Local Execution Package

For FASTQ intake/QC before host-depletion, taxonomic profiling, or functional profiling, use:

```bash
python plugins/ngs-analysis/scripts/run_fastq_assay_package.py \
  --lane shotgun_metagenomics \
  --sample-sheet shotgun_samples.csv \
  --execute
```

This validates read paths and structure, runs seqkit stats and FastQC/MultiQC when available, and writes `taxonomic_classification_status.json`. Add `--kraken-db /path/to/db` only when a local Kraken2 database is available; otherwise the package records the database/tool blocker explicitly.

For backend taxonomic and functional profiling when databases are available, use:

```bash
python plugins/ngs-analysis/scripts/run_shotgun_metagenomics.py \
  --sample-sheet shotgun_samples.csv \
  --kraken-db /db/kraken2/standard \
  --host-reference /refs/human_kneaddata_db \
  --run-bracken \
  --run-humann \
  --humann-db /db/humann \
  --metadata sample_metadata.tsv \
  --execute
```

For nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline taxprofiler`.

When `--host-reference` is supplied, the backend runner adds a KneadData host-depletion step, requires `kneaddata` in tool preflight, writes cleaned FASTQs under `host_depletion/`, and uses those cleaned reads for downstream Kraken2 and HUMAnN steps. Keep the host reference path and host-depletion decision visible because it can change taxonomic and functional abundance conclusions.

The backend runner writes native matrix artifacts when database tools produce outputs:

- `tables/bracken_est_reads_matrix.tsv`
- `tables/bracken_relative_abundance_matrix.tsv`
- `tables/humann_pathabundance_matrix.tsv`
- `tables/humann_genefamilies_matrix.tsv`
- `tables/bracken_summary.json` and `tables/humann_summary.json`
- `tables/top_bracken_taxa.tsv`, `tables/top_humann_pathways.tsv`, `tables/top_humann_gene_families.tsv`, and `tables/metagenomics_backend_review.json` when normalized backend matrices are available

If Kraken2/Bracken/HUMAnN outputs are absent, the summaries and visualization manifest keep those layers `not_available` instead of implying taxonomic or functional interpretation succeeded.

## Kickoff Pattern

nf-core preflight run:

```bash
nextflow run nf-core/taxprofiler \
  -profile test,docker \
  --outdir results/taxprofiler_test
```

Direct Kraken2 skeleton:

```bash
kraken2 \
  --db /path/to/kraken2_db \
  --paired sample_R1.fastq.gz sample_R2.fastq.gz \
  --report results/kraken2/sample.report \
  --output results/kraken2/sample.kraken
```

## Visualization Outputs

The local FASTQ package always writes `visualizations/index.html` and `visualizations/visualization_manifest.json`. With only FASTQs, this is a read-QC/readiness bundle. Provide existing `--kraken-report`, `--bracken-table`, `--humann-pathabundance`, or `--humann-genefamilies` files to generate native taxonomy and functional-profile plots without requiring a Marimo notebook. For full backend runs, `run_shotgun_metagenomics.py` now also merges generated Bracken/HUMAnN outputs into plugin-native tables for the review bundle and writes `visualizations/shotgun_backend_dashboard.html` plus SVG plots for top Bracken taxa, HUMAnN pathways, and HUMAnN gene families when the corresponding matrices are present.

## Guardrails

- Do not auto-download large databases without confirming size and destination.
- Host depletion choices can change biological conclusions; document the reference and parameters.
- Negative controls should stay visible in QC and interpretation.

Referenced files: 1

scrna-seq-qc8.14 KB

View saved version →

---
name: scrna-seq-qc
description: Process, quality-control, annotate, and visualize single-cell or single-nucleus RNA-seq datasets across tissues and species. Use when Codex needs to build, adapt, or review a general scRNA-seq QC pipeline; choose dataset-appropriate cell-level filters from QC distributions; run required scDblFinder-based doublet and ambient-RNA filtering; annotate cells with matched references or marker-based fallbacks; or generate global and per-group UMAP visualizations for large scRNA-seq datasets.
---

# scRNA-seq QC

## Start Here

Read `references/qc-annotation-umap-heuristics.md` before picking thresholds, annotation backends, or UMAP feature-selection rules.

Confirm what inputs exist before writing code:

- An AnnData object or equivalent with raw counts preserved.
- Per-sample, per-batch, or per-channel metadata, because QC and doublet detection should respect technical partitions.
- Organism, tissue, assay type, chemistry, and whether the data are whole-cell or single-nucleus.
- Whether a matched cell atlas or label-transfer reference exists for the tissue and species.

Preserve provenance in the output: package versions, thresholds, threshold-justification plots, counts removed or flagged at each filter, annotation backend and reference, marker-gene selection heuristic, and any manual cluster exclusions.

## Workflow

1. Choose QC thresholds from the data, not from a fixed template.
   - Plot detected genes, total UMIs, mitochondrial fraction, and any tissue-specific nuisance signals overall and by batch.
   - Inspect all available QC metrics, but default filtering should use only the standard metrics: detected genes, total counts, and `percent.mt`.
   - Pick thresholds from the observed distributions and expected biology.
   - Save a plot with threshold lines and record why each threshold is appropriate for this dataset.
   - If another metric looks important enough to filter on, flag it as a dataset-specific issue, explain why, and consult the user before adding that extra filter.
   - Keep QC plots legible: do not overload a single panel with too many batches or categories when faceting, splitting, or summary views would communicate the result more clearly.

2. Run cell-level QC.
   - Remove or flag obvious low-quality barcodes using the chosen thresholds on detected genes, total counts, and `percent.mt`.
   - Use `scDblFinder` for doublet detection. Run it per batch or capture channel, and split very large batches before doublet calling.
   - Do not skip doublet calling or silently substitute another method. If `scDblFinder` cannot run in the environment, surface the blocker explicitly or get user approval before using a different caller.
   - Compute ambient-RNA style metrics and use them for filtering when the dataset and workflow support it.
   - Compute any other informative QC metrics when feasible, but do not turn those additional nonstandard metrics into hard filters without explicit user approval unless the user already asked for a stricter policy.
   - Prefer adding a `passes_QC` column instead of physically dropping cells when downstream provenance matters.

3. Build a latent space and inspect residual artifacts.
   - Decide whether `scVI` is warranted for this dataset and use case before training it.
   - Prefer a standard PCA/Scanpy workflow for smaller, simpler datasets with limited batch structure or when a conventional embedding answers the question cleanly.
   - Prefer `scVI` when integration across batches, donors, chemistries, or related datasets is important, or when the dataset is large and noisy enough that a learned latent space is likely to help.
   - Record why `scVI` or a conventional PCA workflow was chosen for this dataset.
   - Cluster and inspect low-quality, mixed-marker, or ambiguous clusters before downstream visualization.
   - Remove or flag artifact clusters only with explicit evidence, and record the rationale.

4. Annotate cells.
   - If a suitable Allen Brain Cell Atlas reference exists and the dataset is a compatible brain tissue and species, use MapMyCells or `cell_type_mapper`.
   - If no suitable Allen reference exists, use the closest matched reference for tissue, species, assay, and chemistry with an appropriate mapping tool.
   - If no reliable reference exists, annotate conservatively from canonical markers and cluster-level markers. Assign coarse labels first and leave uncertain clusters as unknown or ambiguous rather than overlabeling them.
   - Persist annotation confidence or probability fields when available, together with at least one coarse and one fine label.

5. Choose a general marker panel for global UMAP.
   - Do not rely on a perturbation-specific or brain-only marker panel.
   - Start from HVGs selected in a batch-aware way.
   - Add genes that distinguish major coarse compartments or high-confidence labels, for example top markers per coarse cluster or class.
   - Exclude nuisance-dominated genes if they swamp the embedding unless the biology requires them.
   - Document how the panel was chosen.

6. Generate UMAP visualizations.
   - For a global UMAP, use the learned latent space or the chosen informative marker panel, depending on which better matches the analytical goal and runtime constraints.
   - For per-group UMAPs, subset by a stable coarse label and use the latent representation unless there is a strong reason to rebuild on expression features.
   - Keep plotting separate from filtering so visualization choices do not mutate the core analysis object.
   - Make every plot legible. Use a reasonable number of categories per panel, prefer coarse labels on overview plots, and split or facet figures when fine labels, batches, or neighborhoods would otherwise make the figure unreadable.

7. Scale to large datasets without copying.
   - Keep matrices sparse whenever possible.
   - Avoid densifying whole matrices.
   - Avoid whole-object copies of AnnData or Seurat objects; use views, backed mode, chunked operations, and per-batch or per-group manifests instead.
   - When crossing Python and R boundaries, pass only the subset and metadata required for the step.
   - Write checkpoints after major stages so failures do not require restarting from raw ingest.

## Deliverables

When implementing a pipeline, produce an auditable output set:

- Filtered `.h5ad` or equivalent object with raw counts preserved and QC or annotation fields in metadata.
- QC summary table with input cells, cells removed or flagged by each filter, final cells, and per-batch summaries.
- Threshold-justification plots for detected genes, UMIs, mitochondrial fraction, plus any additional QC metric that was inspected; clearly separate metrics that informed review from metrics that actually drove filtering.
- Parameter manifest with thresholds, package versions, annotation backend and reference, marker-panel heuristic, and any manual exclusions.
- UMAP coordinates and plots for global and per-group views when requested, with category counts and panel layouts chosen so the figures remain legible.

## Embedded Runner

For 10x-style matrix bundles, a local runner is available:

```bash
python plugins/ngs-analysis/scripts/run_scrnaseq_post_count_qc.py --input-dir path/to/scrna_bundle
```

The input directory should contain `matrix/`, `manifest.tsv`, and `dataset_metadata.json`, unless explicit paths are supplied. Treat the runner as an auditable analysis surface: its marker-based fallback is PBMC-oriented when no matched reference is provided, so tissue-specific annotation and integration choices still require review.

The runner writes `visualizations/index.html` for portable artifact review, `summary.md` plus `provenance/analysis_status.json` for explicit completeness/blocker reporting, and auto-launches a localhost Marimo review app recorded in `notebooks/marimo_server.json`. It also writes `notebooks/scrna_qc_review.marimo.py` as a notebook backup over the generated PNG/CSV/H5AD outputs. Treat the notebook and review app as review layers, not as the source of truth; the run envelope and generated artifacts remain canonical.

## Resources

- `references/qc-annotation-umap-heuristics.md`: Threshold-selection heuristics, annotation fallback strategy, general marker-panel selection rules, and large-dataset memory practices.

Referenced files: 2

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
MIT
Package author
OpenAI
Keywords
ngs, sequencing, bioinformatics, fastq, bcl, rnaseq, scrnaseq, variant-calling, atacseq, chipseq, microbiome, metagenomics, pipeline-routing, nextflow, nf-core

Declared capabilities

  • Interactive
  • Read
  • Write

Package observed Sep 30, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 18:00 UTC
Collection status
Collected

Plugin_271fcfe114788191b30908b85bd9ade6

Download plugin data (JSON)