← Files NGS Analysis WorkbenchARCHIVED FILE
references/single-cell-rnaseq.md
5.01 KB · Sep 30, 2026 · 23:20 UTC
# Single-cell RNA-seq guidance Use this guidance for single-cell or single-nucleus RNA-seq starting from FASTQs, count matrices, Cell Ranger-style output, AnnData, Seurat, or related objects. Determine whether the endpoint is count generation, post-count QC, clustering, annotation, visualization, or downstream statistical comparison. Read the basic FASTQ QC reference when raw-read quality interpretation is needed, but let this reference control barcode, UMI, chemistry, and read-role decisions. ## What to inspect - input type and whether raw counts and sample provenance are preserved - single-cell versus single-nucleus assay - chemistry or explicit barcode, UMI, cDNA, index, and whitelist layout - organism, tissue, reference, and expected cells when known - sample, donor, batch, capture channel, and multiplexing metadata - desired endpoint and whether a compatible annotation reference exists Do not infer chemistry from filenames. Do not alter protocol-specific read segments through generic trimming without confirming their biological role. ## FASTQ to count matrix - Choose a counting method compatible with the observed chemistry, read layout, whitelist, organism, and desired output rather than selecting by engine alone. - Treat vendor-standard output as an explicit user choice with its associated terms and constraints, not as the universal default. - Preserve sample and capture identity through barcode and feature matrices. - Record the reference, chemistry, whitelist, cell-calling assumptions, and workflow-generated QC evidence. Counting and post-count interpretation are separate scientific decisions. A successfully generated matrix is not evidence that cell calling, doublets, ambient RNA, sample quality, or annotation are acceptable. ## Post-count QC and annotation - Preserve raw counts and per-sample or per-channel provenance. Make any loss of raw counts or sample identity explicit before continuing. - Choose cell-level thresholds from observed distributions and expected biology, reviewing detected genes, total counts, mitochondrial fraction, and relevant nuisance signals overall and by technical partition. - Treat nonstandard metrics primarily as review signals. Do not turn them into hard filters without dataset-specific evidence and user agreement. - Assess doublets within realistic capture partitions and retain scores and calls. - Assess ambient RNA when supported by the dataset and method. - Prefer a passes-QC annotation over irreversible deletion when auditability matters. - Choose conventional PCA or a learned latent representation according to scale, batch structure, integration needs, and the analytical question. - Match annotation references by species, tissue, assay, chemistry, and biological state. Use conservative hierarchical labels and retain uncertainty. - Keep visualization decisions separate from filtering and preserve legible global and per-group views. - Keep matrices sparse, avoid whole-object copies, and checkpoint major stages for large datasets. Use the detailed single-cell heuristics before choosing specific thresholds, a doublet method, latent-space integration, annotation backend, or UMAP feature panel. ## Replicate-aware downstream comparisons - Treat the biological sample or donor as the experimental unit. Cells and capture channels are not independent biological replicates. - Before differential expression or abundance testing, confirm unique sample and donor identifiers, group assignments, pairing, batch and covariate structure, sufficient biological replication, and an identifiable design. Leave a comparison untested when replication is inadequate or the requested effect is confounded. - For expression comparisons within a cell type or state, preserve raw counts and construct sample- or donor-level pseudobulk counts when that matches the method assumptions. Define the aggregation, inclusion rules, design, baseline, and contrasts explicitly. Do not run an ordinary cell-level test that treats cells from the same sample as independent observations. - Use a replicate-aware cell-level mixed model only when it is deliberately selected for the study design and its assumptions are reviewed and recorded. - For differential abundance, use a replicate-aware method and state the population denominator and sampling assumptions; do not test cell counts as independent events. - Keep post-count QC and annotation, label or state selection, and downstream statistical comparisons as separate reviewed stages. Record which cells, labels, samples, and covariates enter each comparison. ## Evidence to carry forward Preserve the raw-count object, sample metadata, chemistry and reference provenance, cell-calling assumptions, thresholds and justification plots, cells flagged at each decision, doublet and ambient-RNA evidence, latent-space choice, annotation backend and confidence, visualization outputs, biological-replicate identifiers, pseudobulk or mixed-model construction, exact designs and contrasts, package versions, and unresolved biological limitations.
SHA-256: 4f45e84a40e3c0ba33a9307e41702d692544e8f171f9f2aef742fa99147bbe33