{
  "skill_name": "cja-dimension-analysis",
  "evals": [
    {
      "id": 1,
      "prompt": "Analyze the Page Name dimension in my CJA data view. I want to know how many unique pages we have, which pages drive most of our traffic, and whether there are any data quality issues like unspecified or empty values. Generate an HTML report.",
      "expected_output": "Calls findDataViews and setDefaultSessionDataViewId (if needed). Calls findDimensions with 'page name' to resolve the dimension ID (e.g., variables/page), then describeDimension to confirm metadata. Probes last 30 days by default (expands to 90 days if no data). Calls searchDimensionItems(dimensionId, limit: 50000) for cardinality count and classification — Page Name is typically HIGH (1,000–10,000) or VERY_HIGH (>10,000). Calls runReport with dimension rows (≥50 items) and occurrences metric for distribution; computes Gini coefficient, top-1/5/10 % shares, and skew label. Calls searchDimensionItems for bad values ('Unspecified', 'None', '(empty)', 'null', 'undefined'); runs runReport for each matched bad value to count occurrences. Invokes python3 scripts/cja_dimension_analysis.py <json> '<data_view_name>' '<data_view_id>' /tmp --format=html, producing /tmp/dimension_analysis_report_<YYYY-MM-DD_HH-MM>.html containing: executive summary card with dimension name, cardinality badge (e.g., 'HIGH — 4,823 unique values'), skew label, and data quality severity; Chart.js distribution bar chart for top-10 values; cumulative distribution table; error patterns table with count and % per bad value; recommendations panel (performance warning if VERY_HIGH cardinality). Inline chat summary includes: cardinality level + unique count, skew label + top-1 % share, top-10 pages with % shares, missing data % with severity threshold (warning >5%, critical >20%), and Gini coefficient.",
      "files": []
    },
    {
      "id": 2,
      "prompt": "Something looks off with our Marketing Channel dimension over the last 90 days. Can you investigate trends and anomalies — specifically which channels are growing or declining, if any channels appeared or disappeared, and if there were any unusual spikes or drops on specific days?",
      "expected_output": "Calls findDataViews and setDefaultSessionDataViewId (if needed). Calls findDimensions with 'marketing channel' to resolve dimension ID, then describeDimension. Sets date range to last 90 days. Calls runReport with Marketing Channel rows, occurrences metric, and daily/weekly granularity for trend data. Splits 90 days into two 45-day periods; classifies each channel value as Growing (>+10%), Declining (<-10%), Stable, New (present in period 2 only), or Disappeared (present in period 1 only). Applies Z-score anomaly detection (threshold 2.0) against rolling mean/std per channel to flag spikes and drops. Invokes python3 scripts/cja_dimension_analysis.py <json> '<data_view_name>' '<data_view_id>' /tmp --format=html, producing /tmp/dimension_analysis_report_<YYYY-MM-DD_HH-MM>.html containing: trends table with Growing/Declining/Stable/New/Disappeared badge per channel; Chart.js time-series line chart for top channels; anomaly event log table with columns channel, date, type (spike/drop), z-score, and magnitude; period-comparison summary card (period 1 total vs period 2 total). Inline chat summary: each channel with trend badge and % change, any new/disappeared channels with first/last seen dates, anomaly events (channel, date, type, z-score), Z-score threshold used, and 'No anomalies detected' message if none found.",
      "files": []
    },
    {
      "id": 3,
      "prompt": "Compare the Device Type and Browser dimensions side by side, and forecast which values are likely to grow over the next 4 weeks. Use the last 12 weeks of data.",
      "expected_output": "Calls findDataViews and setDefaultSessionDataViewId (if needed). Calls findDimensions twice ('device type' and 'browser') to resolve both dimension IDs, then describeDimension for each. Calls searchDimensionItems(limit: 50000) for cardinality on both (Device Type: LOW/MEDIUM expected; Browser: MEDIUM/HIGH). Calls runReport for distribution (≥50 rows) for each dimension; computes Gini, skew label, top-1/5/10 % shares. Calls searchDimensionItems for bad values for both. Calls runReport with weekly time-series (12 weeks) per dimension for the top 5–10 values; fits linear regression per value, projects 7 weeks forward, computes R², classifies trend direction and confidence. Invokes python3 scripts/cja_dimension_analysis.py <json> '<data_view_name>' '<data_view_id>' /tmp --format=html, producing /tmp/dimension_analysis_report_<YYYY-MM-DD_HH-MM>.html containing: comparison summary card; side-by-side table with rows for Cardinality Level, Unique Values, Skew Label, Gini Coefficient, Top-1 % Share, Missing Data % per dimension column; per-dimension Chart.js distribution bar charts; forecast section with historical + 7-week dashed projected trend lines per top value; confidence badges (High/Medium/Low); recommendations panel. Inline chat summary: comparison table excerpt, top-5 values per dimension with % share, forecast highlights per dimension with direction + confidence + projected end value (e.g., 'Mobile: Upward, High confidence (R²=0.84), projected 58% share in week 16'), flags for high-confidence trends, note if insufficient data for forecasting.",
      "files": []
    }
  ]
}
