← Life Sciences NGS AnalysisCONTENT HISTORY

Update to Life Sciences NGS Analysis

Snapshot Sep 30, 2026 · 22:50 UTC · version 1.0.3

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "ngs-bulk-rnaseq-counts-qc",
  "description": "Run or plan bulk RNA-seq FASTQ-to-count processing with sample-sheet, strandedness, genome annotation, alignment or pseudoalignment, MultiQC, and count-matrix QC checks.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 291
    }
  ],
  "skill_md_contents": "---\nname: ngs-bulk-rnaseq-counts-qc\ndescription: Run or plan bulk RNA-seq FASTQ-to-count processing with sample-sheet, strandedness, genome annotation, alignment or pseudoalignment, MultiQC, and count-matrix QC checks.\n---\n\n# Bulk RNA-seq Counts QC\n\nUse this skill for bulk RNA-seq read processing, quantification, and count-matrix generation. If the user already has a count matrix and wants contrasts or statistics, use `ngs-bulk-rnaseq-differential-expression`.\n\n## Essential Inputs\n\nConfirm:\n\n- FASTQ or aligned-read inputs and paired-end/single-end status\n- organism, genome build, FASTA, GTF, and gene ID convention\n- strandedness or permission to infer strandedness\n- sample sheet with biological condition, replicate, batch, and library metadata\n- desired quantification: gene counts, transcript estimates, or both\n- alignment strategy: `STAR/Salmon`, Salmon-only, featureCounts from BAMs, or existing lab protocol\n\n## Route\n\nPrefer `nf-core/rnaseq` for standard processing when a stable container or HPC runtime is available. Use the `local_light` Snakemake/Salmon path for small local/devbox feasibility runs when Docker, registry egress, or Nextflow process containers are the blocker.\n\nThe plugin-owned local runner is:\n\n```bash\npython plugins/ngs-analysis/scripts/run_bulk_rnaseq_counts_qc.py \\\n  --sample-sheet samplesheet.csv \\\n  --fastq-root path/to/fastqs \\\n  --transcriptome-fasta reference/transcriptome.fasta \\\n  --genome-fasta reference/genome.fa \\\n  --annotation-gtf reference/genes.gtf \\\n  --execute\n```\n\nOmit `--execute` for validation plus Snakemake workflow validation only. Use `--no-dry-run` only when the user wants input validation and run-envelope preparation without workflow graph validation.\n\nThe runner emits a run-local `resources/` readiness bundle with `resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. Resource checks are advisory by default for custom or reduced references; add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when a registered genome bundle must be complete before the run is considered ready.\n\nPreflight command:\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_counts_qc --emit-install-plan\npython plugins/ngs-analysis/scripts/ngs_preflight.py --profile local_light --emit-install-plan\n```\n\n## Decision Points\n\n- If strandedness is unknown, infer it before final counting; do not lock in a design based on library guesses.\n- If strandedness is provided, carry it into the quantification command and flag any disagreement between the configured library type and Salmon's inferred format.\n- Keep genome FASTA, GTF, transcriptome, and aligner indexes from the same build/release.\n- Inspect per-sample reads, mapping rate, rRNA/mitochondrial fraction when available, duplication, insert size, gene-body bias, and assignment rate.\n- Preserve raw counts separately from normalized expression.\n- Carry sample metadata forward exactly; downstream DE depends on this table.\n\n## Outputs\n\nProduce:\n\n- sample sheet and command/profile\n- reference manifest with genome and GTF release\n- MultiQC or equivalent processing summary\n- Salmon `quant.sf` outputs, TPM/NumReads/effective-length matrices, and carried-forward sample metadata\n- Gene-level expected-count and TPM matrices derived from transcript-level Salmon outputs, plus a `tx2gene` provenance table\n- Compact QC verdict JSON covering mapping rate, duplication, library-type agreement, and outlier samples\n- Browser-safe MultiQC helper HTML pages and a localhost launch hint for reliable in-app review\n- Run-local reference readiness artifacts under `resources/`, including the resource plan, manifest, environment exports, and Markdown readiness summary\n- issues that block differential expression, such as missing replicates, mislabeled groups, or severe batch/library failures\n- standard run envelope: `run_manifest.json`, `config.json`, `validation/`, `logs/`, `versions/`, `artifact_index.json`, and `summary.md`\n"
}

SHA-256: 0404efd6ca29e85a8ac35f0acefcaf153aa6af78f9e6494ede6645b2dd28e0f2