{"id":7270,"plugin_id":"Plugin_271fcfe114788191b30908b85bd9ade6","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:50:14.882Z","digest":"866157b164437761b7cb349b6af55884e6e4ed74a30ade4fa160be800f157db0","against":null,"payload":{"name":"ngs-bcl-to-fastq","description":"Validate Illumina BCL run folders and sample sheets, plan demultiplexing, review index/UMI/lane choices, run BCL-to-FASTQ conversion, and interpret demux metrics while surfacing license/download boundaries.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":291}],"skill_md_contents":"---\nname: ngs-bcl-to-fastq\ndescription: Validate Illumina BCL run folders and sample sheets, plan demultiplexing, review index/UMI/lane choices, run BCL-to-FASTQ conversion, and interpret demux metrics while surfacing license/download boundaries.\n---\n\n# BCL To FASTQ\n\nUse this skill when the input is an Illumina BCL run folder or the user asks to demultiplex a sequencing run. This is a deep demultiplexing and run-validation skill, not only a command wrapper.\n\n## Essential Inputs\n\nConfirm:\n\n- run folder path with `RunInfo.xml`\n- sample sheet path and format\n- output directory\n- instrument/run metadata from `RunInfo.xml` and `RunParameters.xml`\n- lane handling: split by lane or combine lanes\n- index mismatch tolerance\n- index read structure and dual-index orientation\n- UMI layout, if any\n- whether adapter trimming/masking should happen during conversion\n- whether undetermined reads and demultiplexing metrics should be reviewed before downstream analysis\n\n## Public Tool Boundary\n\nPrefer `bcl-convert` if it is already installed. It is free for local use but proprietary and RPM-distributed by Illumina, so do not auto-download without explicit user approval.\n\nLegacy `bcl2fastq` may exist in older environments. Use it only when BCL Convert is unavailable or the run requires legacy compatibility.\n\n## Preflight\n\n```bash\npython plugins/ngs-analysis/scripts/ngs_preflight.py --pipeline bcl_to_fastq --emit-install-plan\n```\n\nAlso check run-folder structure:\n\n```bash\ntest -f /path/to/run/RunInfo.xml\ntest -f /path/to/SampleSheet.csv\nfind /path/to/run -maxdepth 4 -type d -name BaseCalls\n```\n\n## Local Execution Package\n\nUse the plugin-owned runner when the user provides a local run folder and sample sheet:\n\n```bash\npython plugins/ngs-analysis/scripts/run_bcl_to_fastq.py \\\n  --run-folder /path/to/run \\\n  --sample-sheet /path/to/SampleSheet.csv \\\n  --output-directory /path/to/fastq_out\n```\n\nAdd `--execute` only when conversion is requested. The runner validates `RunInfo.xml`, optional `RunParameters.xml`, the BaseCalls directory, sample-sheet rows, duplicate lane/index combinations, and index length compatibility. With `--execute`, it uses installed `bcl-convert`, then legacy `bcl2fastq` if available; if neither exists, it records the blocker instead of downloading proprietary software.\n\n## Validation Checklist\n\nBefore conversion, validate:\n\n- `RunInfo.xml` exists and its read structure matches the expected sequencing design.\n- `SampleSheet.csv` exists, is the intended version, and has no duplicate sample/index combinations within each lane.\n- Index sequence lengths match the index reads and any trimming/masking requested by the sample sheet.\n- Dual-index orientation is explicit for the instrument and library prep; do not infer i5 orientation from filenames.\n- UMI bases are assigned to the intended read or index read and carried through to FASTQ headers or output metadata as needed.\n- Lane-splitting, sample-name normalization, and output directory behavior are agreed before running.\n- Disk space is sufficient for output FASTQs, reports, and temporary files.\n\n## Kickoff Pattern\n\nFirst produce a preflight plan with paths and sample sheet validation. Then run conversion only after the user confirms:\n\n```bash\nbcl-convert \\\n  --bcl-input-directory /path/to/run \\\n  --output-directory /path/to/fastq_out \\\n  --sample-sheet /path/to/SampleSheet.csv\n```\n\n## Metrics Review\n\nAfter conversion, inspect and report:\n\n- total clusters, clusters passing filter, and yield by lane\n- percent assigned by sample and percent undetermined by lane\n- top undetermined index sequences when available\n- per-sample FASTQ counts and read-pair consistency\n- unexpected index hopping, barcode collision, or sample-sheet mismatch signals\n\nRecord software version, command, sample sheet checksum, run-folder path, output path, and conversion metrics. Do not start downstream analysis until severe demultiplexing anomalies are surfaced.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}