← Life Sciences NGS AnalysisCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Life Sciences NGS Analysis
Snapshot Sep 30, 2026 · 22:50 UTC · version 1.0.3
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "ngs-fastq-qc",
"description": "Validate FASTQ inputs, run local FastQC/MultiQC QC, interpret QC signals, and optionally execute fastp or Cutadapt trimming branches without overwriting raw reads.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 253
}
],
"skill_md_contents": "---\nname: ngs-fastq-qc\ndescription: Validate FASTQ inputs, run local FastQC/MultiQC QC, interpret QC signals, and optionally execute fastp or Cutadapt trimming branches without overwriting raw reads.\n---\n\n# FASTQ QC\n\nUse this skill for QC-only, trimming-first, or FASTQ quality interpretation workflows. This skill can execute the plugin-owned local FastQ QC runner when the user approves a local run. It should decide whether trimming or additional investigation is warranted; it should not blindly trim by default.\n\n## Essential Inputs\n\nConfirm:\n\n- FASTQ paths and pairing convention\n- whether output should be QC-only or trimmed FASTQs\n- known adapter or primer sequences\n- organism if contamination screening or host depletion is requested\n- output directory\n- whether FASTQs are raw, demultiplexed, previously trimmed, or downloaded from an archive\n- whether downstream analysis expects original read lengths, UMIs, or inline barcodes\n\n## Public Tools\n\nDefault tool set:\n\n- `FastQC` for raw read QC\n- `MultiQC` for project-level summary\n- `fastp` for all-in-one QC/trimming when acceptable\n- `Cutadapt` when primer/adapter handling needs explicit sequences\n- `seqkit` for quick counts, stats, and subsampling\n\n## Preflight\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline fastq_qc --emit-install-plan\n```\n\n## Local Execution\n\nUse the plugin-owned runner for local artifact-producing FASTQ QC:\n\n```bash\npython plugins/ngs-analysis/scripts/run_fastq_qc.py \\\n --sample-sheet samplesheet.csv \\\n --execute\n```\n\nSingle paired sample:\n\n```bash\npython plugins/ngs-analysis/scripts/run_fastq_qc.py \\\n --sample sampleA \\\n --r1 sampleA_R1.fastq.gz \\\n --r2 sampleA_R2.fastq.gz \\\n --execute\n```\n\nOptional trimming branch:\n\n```bash\npython plugins/ngs-analysis/scripts/run_fastq_qc.py \\\n --sample-sheet samplesheet.csv \\\n --trim-mode fastp \\\n --execute\n```\n\nFor explicit adapters:\n\n```bash\npython plugins/ngs-analysis/scripts/run_fastq_qc.py \\\n --sample-sheet samplesheet.csv \\\n --trim-mode cutadapt \\\n --adapter-r1 AGATCGGAAGAGC \\\n --adapter-r2 AGATCGGAAGAGC \\\n --execute\n```\n\nThe runner performs pre-execution validation before Snakemake execution. It writes a timestamped run directory with `run_manifest.json`, `config.json`, `validation/`, `workflow/Snakefile`, logs, `artifact_index.json`, `summary.md`, FastQC/MultiQC outputs, and `qc_interpretation.json` after successful execution.\n\n## Interpretation Rules\n\nInspect raw QC before recommending trimming:\n\n- Per-base quality drop at the read end: consider quality trimming, but preserve enough length for alignment or amplicon merging.\n- Adapter or primer signal: use `cutadapt` when explicit sequences matter; use `fastp` only when automatic handling is acceptable.\n- Poly-G or patterned-flowcell artifacts: handle with a tool that explicitly supports the artifact and report the assumption.\n- Overrepresented sequences: classify adapters, primers, rRNA, PhiX, host contamination, or true biology before filtering.\n- Per-tile failures or severe quality shifts: flag possible run-level issues and avoid treating them as ordinary adapter contamination.\n- High duplication: interpret by assay; it may be expected for amplicons, targeted panels, or low-input libraries.\n- Pairing issues: verify R1/R2 file counts and read-name pairing before any downstream workflow.\n\nDo not overwrite input FASTQs. Preserve the raw QC reports even when trimmed FASTQs are created.\n\n## Kickoff Pattern\n\nQC-only:\n\n```bash\nmkdir -p results/fastqc results/multiqc\nfastqc -t 4 -o results/fastqc *.fastq.gz\nmultiqc results/fastqc -o results/multiqc\n```\n\nQC plus trimming:\n\n```bash\nfastp \\\n -i sample_R1.fastq.gz \\\n -I sample_R2.fastq.gz \\\n -o results/trimmed/sample_R1.fastq.gz \\\n -O results/trimmed/sample_R2.fastq.gz \\\n --html results/fastp/sample.html \\\n --json results/fastp/sample.json\nmultiqc results -o results/multiqc\n```\n\n## Output Review\n\nReturn a short QC interpretation with:\n\n1. sample/read-pair inventory\n2. QC modules that look normal\n3. QC modules that require action or user confirmation\n4. trimming or no-trimming recommendation with rationale\n5. downstream caveats such as short reads, contaminated libraries, or failed pairs\n\nWhen using the local runner, ground the response in the generated `qc_interpretation.json`, `summary.md`, and MultiQC report instead of relying only on expected artifacts.\n"
}SHA-256: b384e8bf977fb1fd04936c1e19c0515442afb9d9d63c93240f0097554a4179ee