{"id":7285,"plugin_id":"Plugin_271fcfe114788191b30908b85bd9ade6","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:50:15.160Z","digest":"a07229c29235fef0ced67075b608d1a1f3769b54f9d393094bae5aebedd542ab","against":null,"payload":{"name":"ngs-shotgun-metagenomics","description":"Kick off public shotgun metagenomics QC, host-depletion, taxonomic profiling, and functional profiling workflows using nf-core/taxprofiler, Kraken2, Bracken, MetaPhlAn, and HUMAnN.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":288}],"skill_md_contents":"---\nname: ngs-shotgun-metagenomics\ndescription: Kick off public shotgun metagenomics QC, host-depletion, taxonomic profiling, and functional profiling workflows using nf-core/taxprofiler, Kraken2, Bracken, MetaPhlAn, and HUMAnN.\n---\n\n# Shotgun Metagenomics\n\nUse this skill for shotgun metagenomic FASTQs.\n\n## Essential Inputs\n\nConfirm:\n\n- paired-end or single-end reads\n- host organism and host-depletion requirement\n- target outputs: taxonomic profile, functional profile, assembly, binning, or QC only\n- preferred database family, if any\n- database paths or permission to download large databases\n- sample metadata, batches, and negative controls\n\n## Public Defaults\n\nPrefer `nf-core/taxprofiler` for reproducible taxonomic profiling. Use direct Kraken2/Bracken, MetaPhlAn, or HUMAnN when the user wants a focused path or already has databases installed.\n\nFor direct backend execution, prefer the plugin runner over handwritten shell when possible because it validates database bundle contents and records `resources/resource_plan.json`, `resource_manifest.tsv`, `resource_env.sh`, and `resource_readiness.md`. `--run-bracken` and `--run-humann` make those database bundles blocking, not merely optional.\n\n## Preflight\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline shotgun_metagenomics --emit-install-plan\n```\n\n## Local Execution Package\n\nFor FASTQ intake/QC before host-depletion, taxonomic profiling, or functional profiling, use:\n\n```bash\npython plugins/ngs-analysis/scripts/run_fastq_assay_package.py \\\n  --lane shotgun_metagenomics \\\n  --sample-sheet shotgun_samples.csv \\\n  --execute\n```\n\nThis validates read paths and structure, runs seqkit stats and FastQC/MultiQC when available, and writes `taxonomic_classification_status.json`. Add `--kraken-db /path/to/db` only when a local Kraken2 database is available; otherwise the package records the database/tool blocker explicitly.\n\nFor backend taxonomic and functional profiling when databases are available, use:\n\n```bash\npython plugins/ngs-analysis/scripts/run_shotgun_metagenomics.py \\\n  --sample-sheet shotgun_samples.csv \\\n  --kraken-db /db/kraken2/standard \\\n  --host-reference /refs/human_kneaddata_db \\\n  --run-bracken \\\n  --run-humann \\\n  --humann-db /db/humann \\\n  --metadata sample_metadata.tsv \\\n  --execute\n```\n\nFor nf-core execution, use `plugins/ngs-analysis/scripts/run_nfcore_pipeline.py --pipeline taxprofiler`.\n\nWhen `--host-reference` is supplied, the backend runner adds a KneadData host-depletion step, requires `kneaddata` in tool preflight, writes cleaned FASTQs under `host_depletion/`, and uses those cleaned reads for downstream Kraken2 and HUMAnN steps. Keep the host reference path and host-depletion decision visible because it can change taxonomic and functional abundance conclusions.\n\nThe backend runner writes native matrix artifacts when database tools produce outputs:\n\n- `tables/bracken_est_reads_matrix.tsv`\n- `tables/bracken_relative_abundance_matrix.tsv`\n- `tables/humann_pathabundance_matrix.tsv`\n- `tables/humann_genefamilies_matrix.tsv`\n- `tables/bracken_summary.json` and `tables/humann_summary.json`\n- `tables/top_bracken_taxa.tsv`, `tables/top_humann_pathways.tsv`, `tables/top_humann_gene_families.tsv`, and `tables/metagenomics_backend_review.json` when normalized backend matrices are available\n\nIf Kraken2/Bracken/HUMAnN outputs are absent, the summaries and visualization manifest keep those layers `not_available` instead of implying taxonomic or functional interpretation succeeded.\n\n## Kickoff Pattern\n\nnf-core preflight run:\n\n```bash\nnextflow run nf-core/taxprofiler \\\n  -profile test,docker \\\n  --outdir results/taxprofiler_test\n```\n\nDirect Kraken2 skeleton:\n\n```bash\nkraken2 \\\n  --db /path/to/kraken2_db \\\n  --paired sample_R1.fastq.gz sample_R2.fastq.gz \\\n  --report results/kraken2/sample.report \\\n  --output results/kraken2/sample.kraken\n```\n\n## Visualization Outputs\n\nThe local FASTQ package always writes `visualizations/index.html` and `visualizations/visualization_manifest.json`. With only FASTQs, this is a read-QC/readiness bundle. Provide existing `--kraken-report`, `--bracken-table`, `--humann-pathabundance`, or `--humann-genefamilies` files to generate native taxonomy and functional-profile plots without requiring a Marimo notebook. For full backend runs, `run_shotgun_metagenomics.py` now also merges generated Bracken/HUMAnN outputs into plugin-native tables for the review bundle and writes `visualizations/shotgun_backend_dashboard.html` plus SVG plots for top Bracken taxa, HUMAnN pathways, and HUMAnN gene families when the corresponding matrices are present.\n\n## Guardrails\n\n- Do not auto-download large databases without confirming size and destination.\n- Host depletion choices can change biological conclusions; document the reference and parameters.\n- Negative controls should stay visible in QC and interpretation.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}