← Life Science ResearchCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Life Science Research
Snapshot Sep 30, 2026 · 23:18 UTC · version 1.0.3
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "locus-to-gene-mapper-skill",
"description": "Map GWAS loci to ranked candidate genes using a deterministic multi-skill chain (EFO -> GWAS -> coordinates -> Open Targets L2G/coloc -> eQTL -> burden/coding context), with reproducible tables and optional figures. Use when a user provides a trait/EFO term and/or lead variants and needs locus-to-gene prioritization for downstream biology decisions.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 113
},
{
"relative_path": "scripts/map_locus_to_gene.py",
"size_in_bytes": 79824
},
{
"relative_path": "scripts/test_map_locus_to_gene.py",
"size_in_bytes": 5903
}
],
"skill_md_contents": "---\nname: locus-to-gene-mapper-skill\ndescription: Map GWAS loci to ranked candidate genes using a deterministic multi-skill chain (EFO -> GWAS -> coordinates -> Open Targets L2G/coloc -> eQTL -> burden/coding context), with reproducible tables and optional figures. Use when a user provides a trait/EFO term and/or lead variants and needs locus-to-gene prioritization for downstream biology decisions.\n---\n\n## Locus-to-Gene Mapper\n\nGenerate a reproducible locus-to-gene mapping for one trait (or a seed set of lead variants), with explicit evidence attribution and conservative confidence labels.\n\nThis skill is optimized for bioinformaticians who need executable, traceable mapping from variant signals to plausible causal genes.\n\n## Required Inputs\n\nProvide at least one anchor source:\n\n- `trait_query` (string), for example `chronic obstructive pulmonary disease`\n- `efo_id` (string), for example `EFO_0000341`\n- `seed_rsids` (list[string]), for example `[\"rs1873625\", \"rs7903146\"]`\n\n## Optional Inputs\n\n- `target_gene` (string), optional gene of interest for highlighting in output\n- `show_child_traits` (bool), default `true`\n- `phenotype_terms` (list[string]), optional additional terms to include when finding anchors\n- `max_anchor_associations` (int), default `1200`\n- `max_loci` (int), default `25`\n- `max_genes_per_locus` (int), default `10`\n- `max_coloc_rows_per_locus` (int), default `100`\n- `max_eqtl_rows_per_variant` (int), default `200`\n- `genebass_burden_sets` (list[string]), default `[\"pLoF\", \"missense|LC\"]`\n- `include_clinvar` (bool), default `true`\n- `include_gnomad_context` (bool), default `true`\n- `include_hpa_tissue_context` (bool), default `true`\n- `include_figures` (bool), default `false`\n- `disable_default_seeds` (bool), default `false`; if `false`, common traits automatically get built-in seed rsIDs\n- `figure_output_dir` (string), default `./output/figures`\n- `mapping_output_path` (string), default `./output/locus_to_gene_mapping.json`\n- `summary_output_path` (string), default `./output/locus_to_gene_summary.md`\n\n## Runtime Requirements\n\n- Python `3.11+`\n- `requests`\n- Optional for figure generation: `matplotlib`, `seaborn`, `pandas`\n\n## Bundled Script (Deterministic Runner)\n\n- Primary entrypoint: `scripts/map_locus_to_gene.py`\n- This script:\n - resolves trait/EFO and anchor variants,\n - resolves seed and anchor rsID coordinates directly through NCBI RefSNP/dbSNP placements,\n - gathers locus-to-gene evidence through the chained skills,\n - writes mapping JSON and summary markdown,\n - optionally renders figures when plotting deps are available.\n\nRun:\n\n```bash\npython locus-to-gene-mapper-skill/scripts/map_locus_to_gene.py \\\n --input-json /path/to/input.json \\\n --print-result\n```\n\nQuick start (no input JSON file):\n\n```bash\npython locus-to-gene-mapper-skill/scripts/map_locus_to_gene.py \\\n --trait-query \"type 2 diabetes\" \\\n --print-result\n```\n\nTrait-only runs default to `include_figures=true` unless explicitly disabled with `--no-include-figures`.\n\nMinimal input JSON:\n\n```json\n{\n \"trait_query\": \"type 2 diabetes\"\n}\n```\n\nBuilt-in default seeds (when `disable_default_seeds=false`):\n\n- `type 2 diabetes` / `t2d` -> `rs7903146`, `rs13266634`, `rs7756992`, `rs5219`, `rs1801282`, `rs4402960`\n- `coronary artery disease` / `cad` -> `rs1333049`, `rs4977574`, `rs9349379`, `rs6725887`, `rs1746048`, `rs3184504`\n- `body mass index` / `bmi` -> `rs9939609`, `rs17782313`, `rs6548238`, `rs10938397`, `rs7498665`, `rs7138803`\n- `asthma` -> `rs7216389`, `rs2305480`, `rs9273349`\n- `rheumatoid arthritis` -> `rs2476601`, `rs3761847`, `rs660895`\n- `alzheimer disease` -> `rs429358`, `rs7412`, `rs6733839`, `rs11136000`, `rs3851179`\n- `ldl cholesterol` / `total cholesterol` -> `rs7412`, `rs429358`, `rs6511720`, `rs629301`, `rs12740374`, `rs11591147`\n\n## Autonomous Execution Contract (Embedded Behavior)\n\nWhen a user asks for locus-to-gene mapping and gives only a trait (for example, `type 2 diabetes`), do the following automatically:\n\n1. Run the bundled script with `--trait-query \"<user_trait>\" --print-result` (no manual JSON required).\n2. If it returns `No anchors remained`, rerun once with a built-in default seed rsID for that trait (unless `disable_default_seeds=true`).\n3. Read the generated `mapping_output_path` and `summary_output_path`.\n4. Return this concise response structure:\n - `Top 5 cross-locus prioritized genes`\n - `Per-locus top gene (score, confidence)`\n - `Visualization artifact` (figure path(s) or Mermaid fallback block)\n - `Warnings and limitations`\n5. For inline image rendering in chat:\n - read `inline_image_markdown` from script result\n - emit those lines exactly as plain markdown (no code fences)\n - if inline rendering still fails, instruct user to upload PNG files into the chat\n\nDo not ask the user to run python manually unless execution is actually blocked.\n\n## Skill Chaining Order (Mandatory)\n\nUse these skills in order. Skip only when an earlier step is not needed by provided inputs.\n\n1. `efo-ontology-skill`\n - Resolve `trait_query` to canonical EFO term and synonyms.\n - Expand descendants when `show_child_traits=true`.\n2. `gwas-catalog-skill`\n - Discover anchor variants for the trait/EFO scope.\n - Pull association/study metadata for locus context.\n3. Built-in NCBI RefSNP coordinate resolution\n - Normalize each anchor rsID to GRCh37/GRCh38 top-level chromosome placements.\n4. `opentargets-skill`\n - Retrieve credible set context, L2G predictions, and colocalisation evidence per locus.\n5. `gtex-eqtl-skill`\n - Retrieve single-tissue eQTL support for anchor variants.\n6. `genebass-gene-burden-skill`\n - Retrieve rare-variant burden support for candidate genes.\n7. `clinvar-variation-skill` (when `include_clinvar=true`)\n - Add variant clinical/coding annotations.\n8. `gnomad-graphql-skill` (when `include_gnomad_context=true`)\n - Add frequency and gene-level constraint context.\n9. `human-protein-atlas-skill` (when `include_hpa_tissue_context=true`)\n - Add tissue plausibility context for top genes.\n\nNever perform additional retrieval after final candidate-gene scoring starts.\n\n## Output Contract (Required)\n\nAlways return:\n\n1. `locus_to_gene_mapping.json`\n2. `locus_to_gene_summary.md`\n\n### JSON contract\n\n```json\n{\n \"meta\": {\n \"trait_query\": \"...\",\n \"efo_id\": \"EFO_...\",\n \"generated_at\": \"ISO-8601\",\n \"sources_queried\": []\n },\n \"anchors\": [\n {\n \"rsid\": \"rs...\",\n \"grch38\": {\"chr\": \"3\", \"pos\": 49629531, \"ref\": \"A\", \"alt\": \"C\"},\n \"lead_trait\": \"...\",\n \"p_value\": 2e-11,\n \"cohort\": \"...\"\n }\n ],\n \"loci\": [\n {\n \"locus_id\": \"chr3:49000000-50200000\",\n \"lead_rsid\": \"rs...\",\n \"candidate_genes\": [\n {\n \"symbol\": \"MST1\",\n \"ensembl_id\": \"ENSG...\",\n \"overall_score\": 0.71,\n \"confidence\": \"High|Medium|Low|VeryLow\",\n \"evidence\": {\n \"l2g_max\": 0.83,\n \"coloc_max_h4\": 0.84,\n \"eqtl_tissues\": [\"Lung\"],\n \"rare_variant_support\": \"none|nominal|strong\",\n \"coding_support\": \"none|noncoding|coding\",\n \"clinvar_support\": \"none|present\",\n \"gnomad_context\": \"...\",\n \"hpa_tissue_support\": [\"lung\"]\n },\n \"rationale\": [\n \"...\"\n ],\n \"limitations\": [\n \"...\"\n ]\n }\n ]\n }\n ],\n \"cross_locus_ranked_genes\": [\n {\n \"symbol\": \"...\",\n \"supporting_loci\": 3,\n \"mean_score\": 0.62,\n \"max_score\": 0.81\n }\n ],\n \"warnings\": [],\n \"limitations\": []\n}\n```\n\n### Markdown summary contract\n\nThe summary must include sections in this exact order:\n\n1. `Objective`\n2. `Inputs and scope`\n3. `Anchor variant summary`\n4. `Per-locus top genes`\n5. `Cross-locus prioritized genes`\n6. `Key caveats`\n7. `Recommended next analyses`\n\n## Optional Figure Contract\n\nOnly produce figures when `include_figures=true`.\n\nIf figures are generated, append this block to JSON:\n\n```json\n{\n \"figures\": [\n {\n \"id\": \"locus_gene_heatmap\",\n \"path\": \"./output/figures/locus_gene_heatmap.png\",\n \"caption\": \"Top candidate genes by evidence component across loci\"\n }\n ]\n}\n```\n\nRecommended figure set:\n\n1. `locus_gene_heatmap.png`\n - Rows: top genes, columns: evidence components (`L2G`, `coloc`, `eQTL`, `burden`, `coding`).\n2. `locus_score_decomposition.png`\n - Stacked bars per locus for top 3 genes.\n3. `tissue_support_dotplot.png`\n - Gene-by-tissue evidence dots from GTEx/HPA context.\n\nIf plotting dependencies are unavailable, skip PNG generation and output Mermaid diagrams in markdown as fallback.\nThe script also returns `inline_image_markdown` and `render_instructions` fields to support inline chat rendering.\n\n## Scoring Rules (Deterministic)\n\nFor each candidate gene per locus, compute:\n\n- `l2g_component`: max L2G score for the gene in locus (`0..1`)\n- `coloc_component`: max `h4` (or `clpp` when only CLPP is available), clipped to `0..1`\n- `eqtl_component`: `min(1, relevant_tissue_hits / 3)`\n- `burden_component`:\n - `1.0` if burden `p < 2.5e-6`\n - `0.6` if `2.5e-6 <= p < 0.05`\n - `0.0` otherwise\n- `coding_component`:\n - `1.0` for coding consequence in target gene with supportive ClinVar annotation\n - `0.6` for coding consequence in target gene without supportive ClinVar annotation\n - `0.3` for noncoding-in-gene support only\n - `0.0` otherwise\n\nOverall score:\n\n`overall_score = 0.40*l2g + 0.25*coloc + 0.15*eqtl + 0.10*burden + 0.10*coding`\n\nConfidence label:\n\n- `High` if score `>= 0.75`\n- `Medium` if `0.55 <= score < 0.75`\n- `Low` if `0.35 <= score < 0.55`\n- `VeryLow` if score `< 0.35`\n\n## Pipeline Contract\n\n### Phase 0: Validate and normalize input\n\n- Enforce that at least one of `trait_query`, `efo_id`, `seed_rsids` is present.\n- Normalize rsID formatting and deduplicate seed variants.\n- Resolve free-text trait to one canonical EFO term when needed.\n\n### Phase 1: Build anchor set\n\n- If trait/EFO input is provided, pull associations and rank anchors by p-value and effect availability.\n- Merge trait-derived anchors with user-supplied `seed_rsids`.\n- Cap anchors using `max_loci` and log dropped anchors in `warnings`.\n\n### Phase 2: Gather locus-to-gene evidence\n\n- Normalize anchor coordinates (both builds when possible).\n- Pull Open Targets locus evidence (credible set/L2G/coloc).\n- Pull GTEx variant-level eQTL rows.\n- Pull gene-level burden results for mapped candidate genes.\n- Pull ClinVar and gnomAD context when enabled.\n\n### Phase 3: Harmonize and score\n\n- Build a per-locus candidate-gene table.\n- Compute deterministic component scores and overall score.\n- Create cross-locus aggregate rankings.\n\n### Phase 4: Synthesize outputs\n\n- Write JSON mapping file.\n- Write markdown summary in exact section order.\n- Optionally generate figures and append `figures` metadata.\n\n### Phase 5: QC gates\n\nFail the run when any of the following occurs:\n\n- No anchors after normalization.\n- Unresolved GRCh38 coordinates should be surfaced as `status=degraded`, not treated as an analytically clean pass.\n- Any locus has candidate genes without score fields.\n- `overall_score` outside `0..1`.\n- Summary section order mismatch.\n- Claim of causality without explicit evidence support in rationale text.\n\n## Public Interface\n\n```python\ndef map_locus_to_gene(input_json: dict) -> dict:\n ...\n```\n\nReturn:\n\n```json\n{\n \"status\": \"ok\",\n \"mapping_output_path\": \"./output/locus_to_gene_mapping.json\",\n \"summary_output_path\": \"./output/locus_to_gene_summary.md\",\n \"figure_paths\": [],\n \"warnings\": [],\n \"limitations\": []\n}\n```\n\n## Non-Invention Rules\n\n- Never invent rsIDs, p-values, scores, cohort labels, tissues, or gene links.\n- Never silently impute missing evidence as positive support.\n- When evidence is missing, record it as a limitation and reduce confidence.\n- Keep evidence provenance explicit (`source skill` + endpoint family) in rationale lines.\n\n## Non-Goals\n\n- Do not claim definitive causal genes from association evidence alone.\n- Do not run fine-mapping methods not directly provided by upstream sources.\n- Do not collapse multiple independent signals into one without stating assumptions.\n"
}SHA-256: 8b8ef7ed433f043236462c224b9827d2a7d30b4b87c8c6dcaf6603f88ae70001