← Life Sciences NGS AnalysisCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Life Sciences NGS Analysis
Snapshot Sep 30, 2026 · 22:50 UTC · version 1.0.3
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Run or plan bulk RNA-seq differential-expression analysis from count matrices with replicate, design formula, contrast, batch, normalization, QC plot, and result-table checks.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 268
}
],
"name": "ngs-bulk-rnaseq-differential-expression",
"skill_md_contents": "---\nname: ngs-bulk-rnaseq-differential-expression\ndescription: Run or plan bulk RNA-seq differential-expression analysis from count matrices with replicate, design formula, contrast, batch, normalization, QC plot, and result-table checks.\n---\n\n# Bulk RNA-seq Differential Expression\n\nUse this skill when the user has raw counts or a count-generation output and wants differential expression, contrasts, QC plots, or ranked gene tables.\n\n## Essential Inputs\n\nConfirm:\n\n- raw count matrix path and sample metadata path\n- gene ID type and annotation mapping requirement\n- biological conditions, replicates, batch variables, donor pairing, covariates, and exclusions\n- exact contrasts and baseline levels\n- preferred statistical framework: DESeq2, edgeR, limma-voom, or existing lab standard\n- output needs: normalized counts, PCA, sample distance, volcano plots, heatmaps, ranked tables, GSEA-ready lists\n\n## Preconditions\n\nDo not start differential expression until:\n\n- raw counts are preserved\n- each requested contrast has enough biological replication\n- sample metadata row names match count matrix columns\n- batch/covariate choices are explicit\n- exploratory PCA/sample-distance plots do not reveal obvious swaps or failed libraries\n\n## Route\n\nFor most count matrices, use DESeq2 or edgeR. Use limma-voom when the study design or lab standard favors it. Keep the analysis in R when using Bioconductor unless the user specifically asks for a Python-only workflow.\n\nThe plugin-owned local runner is:\n\n```bash\npython plugins/ngs-analysis/scripts/run_bulk_rnaseq_de.py \\\n --count-matrix count_matrix.tsv \\\n --sample-metadata sample_metadata.tsv \\\n --contrasts contrasts.tsv \\\n --execute\n```\n\nUse `--method auto` unless the user or lab standard specifies `DESeq2`, `edgeR`, or `limma_log2`. Auto mode uses DESeq2 when integer-like counts and the package are available, falls back to edgeR for integer-like counts, and uses `limma_log2` for non-integer expression matrices.\n\nUse `--input-mode` to declare whether the matrix is `raw_counts`, `normalized_expression`, or `log_expression`. When `--input-mode auto` is used, the runner infers the mode and records a warning if normalization is skipped because the matrix is already transformed.\n\nPreflight command:\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bulk_rnaseq_differential_expression --emit-install-plan\n```\n\n## Decision Points\n\n- Never compare groups without stating the design formula and contrast.\n- Treat batch correction in modeling separately from visual batch removal.\n- Do not filter genes using post-hoc knowledge of the contrast.\n- For paired or repeated-measures designs, model subject/donor explicitly.\n- Report genes with effect size, uncertainty, adjusted p-value, and filtering status.\n\n## Outputs\n\nProduce:\n\n- design formula and contrast manifest\n- QC plots: library size, detected genes, PCA/sample distance, mean-variance trend, and outlier review\n- input-mode-aware matrix exports plus the modeling/log-scale matrix used for DE\n- differential-expression tables per contrast\n- explicit `.not_tested.tsv` stubs for contrasts blocked by insufficient replication or confounding\n- auto-launched localhost Marimo review app recorded in `notebooks/marimo_server.json`\n- caveats for small n, confounded designs, failed samples, or batch variables that cannot be estimated\n- standard run envelope: `run_manifest.json`, `config.json`, `validation/`, `logs/`, `versions/`, `visualizations/`, `notebooks/`, `artifact_index.json`, and `summary.md`\n"
}SHA-256 of public snapshot: a775a42b14ea43047ad556b304ac9d5076c45c973d58c054f7ae879a9213f90e