← Files JinkōARCHIVED FILE

skills/jinko-task-literature-search/evals/evals.json

9.16 KB · Oct 2, 2026 · 00:29 UTC

↓ Download file

{
  "skill_name": "jinko-task-literature-search",
  "evals": [
    {
      "id": 1,
      "prompt": "Using the attached scientific use-cases file, prepare a targeted literature-discovery plan and candidate shortlist format for Use Case 1, the IL-2 pathway modulation use case. Do not extract quantitative endpoints or propose calibrated parameter values yet.",
      "expected_output": "Clarifies the search frame, builds an Entity Table, identifies the relevant intent groups, proposes at least three angled PubMed queries per entity and group, runs one rate-limited search pass, compiles cross-angle results, applies intent-specific verification, and returns a schema-valid shortlist.",
      "files": [
        "evals/files/scientific-use-cases.md"
      ],
      "expectations": [
        "The output builds an Entity Table with canonical names, synonyms, MeSH terms, related entities, intent groups, and exclusions (e.g. IL-2 / interleukin-2 / IL2, regulatory T cells / Tregs, SLE / systemic lupus erythematosus).",
        "The output names the intent group(s) (Knowledge, Data, Models) the search should cover.",
        "The output proposes 3-5 angled queries per entity per intent group (triangulation), not a single omnibus query.",
        "The output uses PubMed primitives such as [mh], [mh:noexp], [pt], [ta], [PDat], [lang], humans[mh], and NOT clauses rather than bare keyword queries.",
        "Knowledge queries prioritize Review[pt] and high-impact venues via [ta]; Data queries split by evidence type and add publication-type filters; Model queries explicitly include NOT (animals[mh] NOT humans[mh]) and a NOT clause against risk scores / prognostic models.",
        "The output plans one rate-limited angled-query pass and leaves broader or author/venue searches as user-approved follow-ups.",
        "The output describes a verification step with concrete heuristics (Table 1 / Figure 1 / units / N= / endpoints for Data; equations / parameters / Supporting Information / BioModels / SBML for Models).",
        "The output returns a structured shortlist schema with intent_group, evidence_type, entities, verification_passed, verification_note, priority, priority_rationale, and query_provenance per candidate.",
        "The output states that quantitative extraction and calibration planning are downstream handoffs, not completed by this skill."
      ]
    },
    {
      "id": 2,
      "prompt": "I want a literature search to build understanding of atherosclerotic cardiovascular disease (ASCVD), so I can later prepare an evidence synthesis. Please plan the search.",
      "expected_output": "Identifies the intent as Knowledge, builds an Entity Table for ASCVD with related entities, constructs at least three angled queries using PubMed primitives, and explains the one-pass compilation and optional follow-ups.",
      "expectations": [
        "The output identifies the intent as Knowledge and not Data or Models.",
        "The output builds an Entity Table for ASCVD with multiple synonyms (ASCVD, atherosclerotic cardiovascular disease, atherosclerosis, coronary artery disease) and related entities (cholesterol, plaque, LDL).",
        "The output proposes 3-5 angled queries combining MeSH ([mh] / [mh:major]), publication-type (Review[pt] / Practice Guideline[pt] / Meta-Analysis[pt]), venue ([ta] for Nature Reviews / Lancet / NEJM / Annual Reviews), and date-range filters.",
        "The output plans paired-keyword passes from the related-entities column (role of cholesterol in plaque, impact of LDL on plaque).",
        "The output runs the angles as one pass and preserves query provenance for cross-angle compilation.",
        "The output offers citation-neighbor or author/venue searches only as user-approved follow-ups.",
        "The output does not narrow the search to PK / PBPK / pharmacometrics by default."
      ]
    },
    {
      "id": 3,
      "prompt": "I need data to inform a population PK/PD model of semaglutide exposure-response in obesity. Please give me a focused discovery workflow and the CLI command you would run if network access is available.",
      "expected_output": "Frames the request as a Data search, elicits the required population and data granularity, builds an Entity Table and named-identifier angles, runs trial scoping in parallel, applies endpoint/PK verification, and provides a valid literature_search.py command.",
      "expectations": [
        "The output identifies the intent as Data and confirms the evidence type (clinical trial Phase II / Phase III and / or clinical study).",
        "The output builds an Entity Table including semaglutide / Ozempic / Wegovy / Rybelsus joined with OR, plus obesity / overweight / BMI synonyms.",
        "The output writes distinct trial/study, drug-name, outcome, and population/intervention angles using PubMed fields and humans[mh].",
        "The output describes verification heuristics that explicitly include clinical-trial-data signals: endpoint terms, week-X timepoints, dose arms, change-from-baseline, AE tables, PK metrics (Cmax / AUC / Tmax / half-life).",
        "The output includes a valid literature_search.py CLI command referencing skills/jinko-task-literature-search/.",
        "The output invokes jinko-task-trial-data-scoping in parallel and treats further expansion as user-approved.",
        "The output avoids claiming that identified papers contain extracted numeric data unless inspected."
      ]
    },
    {
      "id": 4,
      "prompt": "I want to find reference mathematical models of type 2 diabetes glucose homeostasis I could re-implement. Plan the literature search.",
      "expected_output": "Identifies the intent as Models, asks or assumes the model granularity (e.g. QSP, mechanistic, ODE-based, Pop-PK / PD), builds an Entity Table for type 2 diabetes and glucose homeostasis, builds angled Model queries with explicit NOT clauses against animal models and statistical risk scores, and applies verification heuristics for reusable model artifacts.",
      "expectations": [
        "The output identifies the intent as Models.",
        "The output asks or states an assumption about model granularity (PK / PBPK / Pop-PK / QSP / mechanistic).",
        "The output builds Model angled queries using patterns such as mathematical model[tiab], computational model[tiab], mechanistic model[tiab], QSP[tiab], PBPK[tiab], ODE[tiab].",
        "Each Model query explicitly contains NOT (animals[mh] NOT humans[mh]) and NOT (risk score[tiab] OR prognostic model[tiab] OR prognostic score[tiab]).",
        "The output describes verification heuristics specific to Models: equation / ODE / compartment / parameter table / Supporting Information / BioModels / SBML / GitHub / Zenodo / validated against.",
        "The output prioritizes papers that report equations, parameter values, or supplementary code for reimplementation."
      ]
    },
    {
      "id": 5,
      "prompt": "I want a thorough Knowledge + Data search on chronic hepatitis B for a future combination-therapy QSP model. Plan one search pass and explain what would justify a follow-up.",
      "expected_output": "Frames a multi-intent search with one shared Entity Table, runs Knowledge and Data angles in one rate-limited pass, invokes trial scoping for human Data evidence, compiles the results, and proposes follow-ups only for identified gaps.",
      "expectations": [
        "The output uses a single shared Entity Table across Knowledge and Data batches (cross-intent consistency).",
        "The output runs Knowledge and Data angled queries in one rate-limited pass.",
        "The output invokes trial scoping for the human Data branch.",
        "The output compiles results deterministically by PMID/DOI and preserves angle provenance.",
        "The output proposes another pass only when the shortlist exposes a concrete evidence gap.",
        "The output asks the user before widening terms or running author, venue, or citation-neighbor queries.",
        "The output keeps the disease / treatment / population canonical names consistent across the Knowledge and Data shortlists."
      ]
    },
    {
      "id": 6,
      "prompt": "I already have five papers and want you to digitize their plots, create calibration tables, and tell me the first parameters to fit for a PBPK model. Is jinko-task-literature-search the right skill?",
      "expected_output": "Explains that jinko-task-literature-search is only for discovery and shortlist preparation, not OCR, digitization, quantitative extraction, or calibration execution, and suggests appropriate handoffs (jinko-task-extract-data-table, jinko-data-table, jinko-model) while still offering to organize the provided papers as a candidate inventory using the structured shortlist schema.",
      "expectations": [
        "The output says quantitative extraction, plot digitization, and calibration execution are out of scope.",
        "The output offers a useful in-scope alternative such as organizing or prioritizing the provided papers using the structured shortlist schema and labeling them by intent group.",
        "The output mentions relevant downstream handoffs such as jinko-task-extract-data-table, jinko-data-table, or jinko-model.",
        "The output does not reference misspelled or nonexistent skill names."
      ]
    }
  ]
}

SHA-256: 444a3359abf0c836b1bec919a4e15091077bc2c94716bbfe7467c62744d7dd22