← LaunchDarklyCONTENT HISTORY

Update to LaunchDarkly

Snapshot Sep 30, 2026 · 23:09 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "online-evals",
  "description": "Attach judges to config variations for automatic LLM-as-a-judge evaluation. Create custom judges, configure sampling rates, and monitor quality scores.",
  "included_files": [
    {
      "relative_path": "README.md",
      "size_in_bytes": 1290
    }
  ],
  "skill_md_contents": "---\nname: online-evals\ndescription: Attach judges to config variations for automatic LLM-as-a-judge evaluation. Create custom judges, configure sampling rates, and monitor quality scores.\ncompatibility: Requires LaunchDarkly API access token with ai-configs:write permission. SDK versions Python v0.20.0+ or Node.js v0.20.0+ for automatic metric recording and the consolidated `track_judge_result` / `trackJudgeResult` API.\nmetadata:\n  author: launchdarkly\n  version: \"0.1.0\"\n---\n\n# Config Online Evaluations\n\nAttach judges to config variations for automatic quality scoring using LLM-as-a-judge methodology. Judges evaluate responses and return scores between 0.0 and 1.0.\n\n## Prerequisites\n\n- LaunchDarkly account with AgentControl enabled\n- API access token with write permissions\n- Existing config with variations (use `configs-create` skill)\n- For automatic metric recording and the consolidated judge-result API: Python AI SDK v0.20.0+ or Node.js AI SDK v0.20.0+\n\n## API Key Detection\n\n1. **Check environment variables** - `LAUNCHDARKLY_API_KEY`, `LAUNCHDARKLY_API_TOKEN`, `LD_API_KEY`\n2. **Check MCP config** - Claude: `~/.claude/config.json` -> `mcpServers.launchdarkly.env.LAUNCHDARKLY_API_KEY`\n3. **Prompt user** - Only if detection fails\n\n## Core Concepts\n\n### What Are Judges?\n\nJudges are specialized configs in **judge mode** that evaluate responses from other configs. They use an LLM to score outputs and return structured results:\n\n```json\n{\n  \"score\": 0.85,\n  \"reasoning\": \"Answered correctly with one minor omission\"\n}\n```\n\n### Built-in Judges\n\nLaunchDarkly provides three pre-configured judges:\n\n| Judge | Metric Key | Measures |\n|-------|-----------|----------|\n| Accuracy | `$ld:ai:judge:accuracy` | How correct and grounded the response is |\n| Relevance | `$ld:ai:judge:relevance` | How well it addresses the user request |\n| Toxicity | `$ld:ai:judge:toxicity` | Harmful or unsafe phrasing (lower = safer) |\n\n### Completion Mode Only\n\nJudges can only be attached to **completion mode** configs in the UI. For agent mode or custom pipelines, use programmatic evaluation via the SDK.\n\n### Restrictions\n\n- Cannot attach judges to judges (no recursion)\n- Cannot attach multiple judges with the same metric key to a single variation\n- Cannot view/edit model parameters or tools on judge variations\n\n## Workflow\n\n### Step 1: Create Custom Judges (Optional)\n\nFor domain-specific evaluation, create judge configs:\n\n```bash\n# Create judge config\ncurl -X POST \"https://app.launchdarkly.com/api/v2/projects/{projectKey}/ai-configs\" \\\n  -H \"Authorization: {api_token}\" \\\n  -H \"Content-Type: application/json\" \\\n  -H \"LD-API-Version: beta\" \\\n  -d '{\n    \"key\": \"security-judge\",\n    \"name\": \"Security Judge\",\n    \"mode\": \"judge\",\n    \"evaluationMetricKey\": \"security\",\n    \"isInverted\": false\n  }'\n```\n\n> **Note:** Set `isInverted: true` for metrics like toxicity where 0.0 is better.\n\nThen add a variation with the evaluation prompt:\n\n```bash\ncurl -X POST \"https://app.launchdarkly.com/api/v2/projects/{projectKey}/ai-configs/security-judge/variations\" \\\n  -H \"Authorization: {api_token}\" \\\n  -H \"Content-Type: application/json\" \\\n  -H \"LD-API-Version: beta\" \\\n  -d '{\n    \"key\": \"default\",\n    \"name\": \"Default\",\n    \"messages\": [\n      {\n        \"role\": \"system\",\n        \"content\": \"You are a security auditor. Score from 0.0 to 1.0:\\n- 1.0: No security issues\\n- 0.7-0.9: Minor issues\\n- 0.4-0.6: Moderate issues\\n- 0.1-0.3: Serious vulnerabilities\\n- 0.0: Critical vulnerabilities\\n\\nCheck for: SQL injection, XSS, hardcoded secrets, command injection.\"\n      }\n    ],\n    \"modelConfigKey\": \"OpenAI.gpt-4o-mini\",\n    \"model\": {\n      \"parameters\": {\n        \"temperature\": 0.3\n      }\n    }\n  }'\n```\n\n### Step 2: Attach Judges to Variations\n\nUse the variation PATCH endpoint:\n\n```bash\ncurl -X PATCH \"https://app.launchdarkly.com/api/v2/projects/{projectKey}/ai-configs/{configKey}/variations/{variationKey}\" \\\n  -H \"Authorization: {api_token}\" \\\n  -H \"Content-Type: application/json\" \\\n  -H \"LD-API-Version: beta\" \\\n  -d '{\n    \"judgeConfiguration\": {\n      \"judges\": [\n        {\"judgeConfigKey\": \"security-judge\", \"samplingRate\": 1.0},\n        {\"judgeConfigKey\": \"api-contract-judge\", \"samplingRate\": 0.5}\n      ]\n    }\n  }'\n```\n\n> **Important:** The `judges` array **replaces all existing** judge attachments. An empty array removes all judges.\n\n### Step 3: Set Fallthrough on Judges\n\nEach judge config needs its fallthrough set to the enabled variation. Configs default to the \"disabled\" variation (index 0).\n\n> **Note:** `turnTargetingOn` does not work for configs. Use `updateFallthroughVariationOrRollout` instead.\n\n```bash\n# First get the variation ID for \"Default\" from GET targeting response\ncurl -X PATCH \"https://app.launchdarkly.com/api/v2/projects/{projectKey}/ai-configs/security-judge/targeting\" \\\n  -H \"Authorization: {api_token}\" \\\n  -H \"Content-Type: application/json; domain-model=launchdarkly.semanticpatch\" \\\n  -H \"LD-API-Version: beta\" \\\n  -d '{\n    \"environmentKey\": \"production\",\n    \"instructions\": [{\n      \"kind\": \"updateFallthroughVariationOrRollout\",\n      \"variationId\": \"your-default-variation-uuid\"\n    }]\n  }'\n```\n\n## Python Implementation\n\n```python\nimport requests\nimport os\nfrom typing import Optional\n\nclass AIConfigJudges:\n    \"\"\"Manager for config judge attachments\"\"\"\n\n    def __init__(self, api_token: str, project_key: str):\n        self.api_token = api_token\n        self.project_key = project_key\n        self.base_url = \"https://app.launchdarkly.com/api/v2\"\n        self.headers = {\n            \"Authorization\": api_token,\n            \"Content-Type\": \"application/json\",\n            \"LD-API-Version\": \"beta\"\n        }\n\n    def attach_judges(self, config_key: str, variation_key: str,\n                      judges: list[dict]) -> dict:\n        \"\"\"\n        Attach judges to a variation.\n\n        Args:\n            config_key: config key\n            variation_key: Variation key\n            judges: List of {\"judgeConfigKey\": str, \"samplingRate\": float}\n        \"\"\"\n        url = f\"{self.base_url}/projects/{self.project_key}/ai-configs/{config_key}/variations/{variation_key}\"\n\n        response = requests.patch(url, headers=self.headers, json={\n            \"judgeConfiguration\": {\"judges\": judges}\n        })\n\n        if response.status_code == 200:\n            print(f\"[OK] Attached {len(judges)} judges to {config_key}/{variation_key}\")\n            return response.json()\n        print(f\"[ERROR] {response.status_code}: {response.text}\")\n        return {}\n\n    def create_judge(self, key: str, name: str, metric_key: str,\n                     system_prompt: str, model: str = \"OpenAI.gpt-4o-mini\",\n                     is_inverted: bool = False) -> dict:\n        \"\"\"\n        Create a judge config.\n\n        Args:\n            key: Judge config key\n            name: Display name\n            metric_key: Metric key for scoring (appears as $ld:ai:judge:{metric_key})\n            system_prompt: Evaluation instructions\n            is_inverted: True if lower scores are better (e.g., toxicity)\n        \"\"\"\n        # Create config\n        config_url = f\"{self.base_url}/projects/{self.project_key}/ai-configs\"\n        response = requests.post(config_url, headers=self.headers, json={\n            \"key\": key,\n            \"name\": name,\n            \"mode\": \"judge\",\n            \"evaluationMetricKey\": metric_key,\n            \"isInverted\": is_inverted\n        })\n\n        if response.status_code not in [200, 201]:\n            print(f\"[ERROR] Creating config: {response.text}\")\n            return {}\n\n        # Create variation\n        var_url = f\"{self.base_url}/projects/{self.project_key}/ai-configs/{key}/variations\"\n        response = requests.post(var_url, headers=self.headers, json={\n            \"key\": \"default\",\n            \"name\": \"Default\",\n            \"messages\": [{\"role\": \"system\", \"content\": system_prompt}],\n            \"modelConfigKey\": model,\n            \"model\": {\"parameters\": {\"temperature\": 0.3}}\n        })\n\n        if response.status_code in [200, 201]:\n            print(f\"[OK] Created judge: {key}\")\n            return response.json()\n        print(f\"[ERROR] Creating variation: {response.text}\")\n        return {}\n\n    def set_fallthrough(self, config_key: str, environment: str,\n                        variation_key: str = \"default\") -> bool:\n        \"\"\"\n        Set fallthrough to enable a judge config.\n\n        Note: turnTargetingOn doesn't work for configs. Instead, set the\n        fallthrough from disabled (index 0) to the enabled variation.\n        \"\"\"\n        # Get variation ID\n        url = f\"{self.base_url}/projects/{self.project_key}/ai-configs/{config_key}/targeting\"\n        response = requests.get(url, headers=self.headers)\n\n        if response.status_code != 200:\n            print(f\"[ERROR] {response.status_code}: {response.text}\")\n            return False\n\n        targeting = response.json()\n        variation_id = None\n        for var in targeting.get(\"variations\", []):\n            if var.get(\"key\") == variation_key or var.get(\"name\") == variation_key:\n                variation_id = var.get(\"_id\")\n                break\n\n        if not variation_id:\n            print(f\"[ERROR] Variation '{variation_key}' not found\")\n            return False\n\n        # Set fallthrough\n        response = requests.patch(url, headers={\n            **self.headers,\n            \"Content-Type\": \"application/json; domain-model=launchdarkly.semanticpatch\"\n        }, json={\n            \"environmentKey\": environment,\n            \"instructions\": [{\n                \"kind\": \"updateFallthroughVariationOrRollout\",\n                \"variationId\": variation_id\n            }]\n        })\n\n        if response.status_code == 200:\n            print(f\"[OK] Fallthrough set for {config_key}\")\n            return True\n        print(f\"[ERROR] {response.status_code}: {response.text}\")\n        return False\n```\n\n## SDK: Automatic Evaluation\n\nWhen using `create_model()` + `run()`, attached judges evaluate automatically:\n\n```python\nimport os\nimport json\nimport asyncio\nimport ldclient\nfrom ldclient import Context\nfrom ldclient.config import Config\nfrom ldai import LDAIClient, AICompletionConfigDefault\n\nsdk_key = os.getenv('LAUNCHDARKLY_SDK_KEY')\nai_config_key = os.getenv('LAUNCHDARKLY_AI_CONFIG_KEY', 'sample-ai-config')\n\nasync def async_main():\n    ldclient.set_config(Config(sdk_key))\n    aiclient = LDAIClient(ldclient.get())\n\n    context = (\n        Context.builder('example-user-key')\n        .kind('user')\n        .name('Sandy')\n        .build()\n    )\n\n    default_value = AICompletionConfigDefault(enabled=False)\n\n    # create_model() initializes with judges from Config\n    model = await aiclient.create_model(ai_config_key, context, default_value, {})\n\n    if not model:\n        print(f\"agent configuration not enabled for: {ai_config_key}\")\n        return\n\n    user_input = 'How can LaunchDarkly help me?'\n\n    # run() automatically evaluates with attached judges\n    result = await model.run(user_input)\n    print(\"Response:\", result.content)\n\n    # Await evaluation results\n    if result.evaluations and len(result.evaluations) > 0:\n        eval_results = await asyncio.gather(*result.evaluations)\n        results_to_display = [\n            r.to_dict() if r is not None else \"not evaluated\"\n            for r in eval_results\n        ]\n        print(\"Judge results:\")\n        print(json.dumps(results_to_display, indent=2, default=str))\n\n    # Always flush events before closing — trailing events are at risk of being\n    # lost otherwise, in short-lived scripts and long-running services alike.\n    ldclient.get().flush()\n    ldclient.get().close()\n```\n\n## SDK: Direct Judge Evaluation\n\nFor agent mode or custom pipelines, evaluate input/output pairs directly:\n\n```python\nimport os\nimport json\nimport asyncio\nimport ldclient\nfrom ldclient import Context\nfrom ldclient.config import Config\nfrom ldai import LDAIClient, AIJudgeConfigDefault\n\nsdk_key = os.getenv('LAUNCHDARKLY_SDK_KEY')\njudge_key = os.getenv('LAUNCHDARKLY_AI_JUDGE_KEY', 'sample-ai-judge-accuracy')\n\nasync def async_main():\n    ldclient.set_config(Config(sdk_key))\n    aiclient = LDAIClient(ldclient.get())\n\n    context = (\n        Context.builder('example-user-key')\n        .kind('user')\n        .name('Sandy')\n        .build()\n    )\n\n    judge_default_value = AIJudgeConfigDefault(enabled=False)\n\n    # Get judge configuration from LaunchDarkly\n    judge = aiclient.create_judge(judge_key, context, judge_default_value)\n\n    if not judge:\n        print(f\"agent judge configuration not enabled for key: {judge_key}\")\n        return\n\n    input_text = 'You are a helpful assistant. How can you help me?'\n    output_text = 'I can answer any question you have.'\n\n    # Evaluate the input/output pair — returns a JudgeResult.\n    judge_result = await judge.evaluate(input_text, output_text)\n\n    if not judge_result.sampled:\n        print(\"Judge evaluation was skipped (sample rate or configuration issue)\")\n        return\n\n    # Track the consolidated result on the Config tracker if needed:\n    # tracker = ai_config.create_tracker()\n    # tracker.track_judge_result(judge_result)\n\n    print(\"Judge Result:\")\n    print(json.dumps(judge_result.to_dict(), default=str))\n\n    # Always flush events before closing — trailing events are at risk of being\n    # lost otherwise, in short-lived scripts and long-running services alike.\n    ldclient.get().flush()\n    ldclient.get().close()\n```\n\n> **Note:** Direct evaluation does not automatically record metrics. Obtain a tracker via `ai_config.create_tracker()` / `aiConfig.createTracker()` and call `tracker.track_judge_result(result)` / `tracker.trackJudgeResult(result)` to record scores for the config you're evaluating.\n\n## Sampling Rates\n\nEach evaluated response sends an additional request to your model provider, increasing token usage and costs. Start with a lower sampling percentage and increase only if you need more evaluation coverage.\n\nYou can adjust sampling rates at any time from the Judges section of a variation, or disable a judge by setting its sampling to 0%.\n\n## Viewing Results\n\n1. Navigate to **configs** > select your config\n2. Click **Monitoring** tab\n3. Select **Evaluator metrics** from dropdown\n4. View scores by variation and time range\n\nResults appear within 1-2 minutes of evaluation.\n\n## Use in Guardrails and Experiments\n\nEvaluation metrics integrate with:\n- **Guarded rollouts**: Pause/revert when scores fall below threshold\n- **Experiments**: Compare variations using evaluation metrics as goals\n\n## Error Handling\n\n| Status | Cause | Solution |\n|--------|-------|----------|\n| 404 | Config/variation not found | Verify keys exist |\n| 400 | Invalid judge config | Check judgeConfigKey exists |\n| 403 | Insufficient permissions | Check API token permissions |\n| 422 | Duplicate metric key | Cannot attach multiple judges with same metric key |\n\n## Next Steps\n\nAfter attaching judges:\n1. **Set fallthrough** on judge configs to an enabled variation (required)\n2. **Monitor results** in Monitoring tab\n3. **Adjust sampling** based on cost/coverage needs\n4. **Set up guarded rollouts** for automatic regression detection\n\n## Related Skills\n\n- `configs-create` - Create configs and judges\n- `configs-targeting` - Configure targeting rules\n- `configs-variations` - Manage variations\n\n## References\n\n- [Online Evaluations](https://docs.launchdarkly.com/home/ai-configs/online-evaluations.md)\n- [Custom Judges](https://docs.launchdarkly.com/home/ai-configs/custom-judges.md)\n\n**Python SDK examples:**\n- [create_judge_example.py](https://github.com/launchdarkly/hello-python-ai/blob/main/features/create_judge/create_judge_example.py) - Evaluate input/output pairs directly via `create_judge` + `evaluate`\n- [create_model_example.py](https://github.com/launchdarkly/hello-python-ai/blob/main/features/create_model/create_model_example.py) - Automatic evaluation with `create_model` + `run` (attached judges fire during the run)\n\n**Node.js SDK examples:**\n- [features/create-judge](https://github.com/launchdarkly/js-core/blob/main/packages/sdk/server-ai/examples/features/create-judge/src/index.ts) - Evaluate input/output pairs directly via `createJudge` + `evaluate`\n- [features/create-model](https://github.com/launchdarkly/js-core/blob/main/packages/sdk/server-ai/examples/features/create-model/src/index.ts) - Automatic evaluation with `createModel` + `run` (attached judges fire during the run)\n"
}

SHA-256: f360edcf428cc59311743d2c391623d54e54a163f307de499e357e6be13cabba