← Life Sciences NGS AnalysisCONTENT HISTORY

Update to Life Sciences NGS Analysis

Snapshot Sep 30, 2026 · 22:50 UTC · version 1.0.3

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "ngs-dna-germline-variants",
  "description": "Run or plan deep germline WGS, WES, targeted-panel, cohort, or trio variant-calling workflows with reference-build, known-sites, QC, joint-calling, and annotation checks.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 300
    }
  ],
  "skill_md_contents": "---\nname: ngs-dna-germline-variants\ndescription: Run or plan deep germline WGS, WES, targeted-panel, cohort, or trio variant-calling workflows with reference-build, known-sites, QC, joint-calling, and annotation checks.\n---\n\n# Germline DNA Variants\n\nUse this skill for germline WGS, WES, or inherited-disease panel analysis from FASTQ, BAM, or CRAM. If the request is tumor-only, tumor-normal, or low-frequency molecular-barcode panel calling, use a somatic or UMI-panel skill instead.\n\n## Essential Inputs\n\nConfirm:\n\n- data type: WGS, WES, or targeted panel\n- sample model: singleton, cohort, duo, trio, family, or case/control\n- input type: FASTQ, BAM, or CRAM\n- organism, reference build, FASTA, indexes, and contig naming\n- known-sites resources for BQSR, contamination, and annotation\n- target BED and bait BED for WES/panel data\n- sex/ploidy assumptions and mitochondrial/sex-chromosome requirements\n- desired callers, annotation outputs, and final VCF/gVCF expectations\n\n## Route\n\nPrefer `nf-core/sarek` for full FASTQ/BAM-to-VCF workflows. Use direct GATK4, DeepVariant, samtools, or bcftools only for focused tasks or a custom workflow.\n\nPreflight command:\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline dna_germline_variants --emit-install-plan\n```\n\nFor compact local checks from prepared BAM/CRAM files, use the shared DNA execution package:\n\n```bash\npython plugins/ngs-analysis/scripts/run_dna_variant_calling.py \\\n  --sample-sheet dna_samples.tsv \\\n  --reference-fasta reference.fa \\\n  --execute\n```\n\nTreat this as a focused samtools/bcftools run envelope, not as a substitute for full cohort, trio, gVCF, BQSR, or annotation workflows.\n\nFor a higher-fidelity local germline run that owns BQSR, per-sample gVCFs, and joint genotyping assumptions, use the germline-specific runner:\n\n```bash\npython plugins/ngs-analysis/scripts/run_dna_germline_variants.py \\\n  --sample-sheet dna_samples.tsv \\\n  --reference-fasta reference.fa \\\n  --known-sites dbsnp.vcf.gz \\\n  --known-sites mills.vcf.gz \\\n  --emit-gvcf \\\n  --joint-call \\\n  --execute\n```\n\nThis runner still expects reference-matched resources and an available GATK toolchain. It packages the validation state and generated artifacts even when execution is blocked by missing tools or resources.\n\nIt also writes advisory `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md` artifacts by default. Add `--genome-build`, `--bundle-root <bundle>=<path>`, and `--require-resource-plan` when complete registered reference and known-sites bundles should be mandatory for readiness.\n\n## Decision Points\n\n- For cohorts or families, decide whether the endpoint is per-sample VCFs, gVCFs for joint genotyping, or a jointly called cohort VCF.\n- For WES/panels, carry the target BED through alignment metrics, calling, and coverage reports; do not call off-target regions by accident.\n- Use BQSR only when reference-matched known-sites resources exist. Do not mix GRCh37, hg19, GRCh38, or T2T resources.\n- Check sample identity, sex concordance, contamination, coverage, duplication, insert size, and transition/transversion where feasible.\n- For trios, preserve pedigree metadata and report Mendelian/QC checks separately from variant interpretation.\n\n## Outputs\n\nProduce:\n\n- command or workflow profile and sample sheet\n- reference/resource manifest with versions and checksums when available\n- QC summary: coverage, duplication, insert size, contamination, sex/relatedness checks when run\n- VCF/gVCF path, index path, and annotation path\n- limitations: low coverage, missing known-sites, target design gaps, or build mismatches\n\nClinical interpretation, pathogenicity classification, and report signing are out of scope unless the user provides a validated clinical workflow.\n"
}

SHA-256: ec855ba0fdfe402b5d9c095d235d80a78a68bf67d5ccfc2b9e929457b9e9c94f