← Files RAG & GenAI CopilotARCHIVED FILE

skills/production-rag-genai-copilot/references/rag_evaluation_framework.md

3.12 KB · Sep 30, 2026 · 23:18 UTC

↓ Download file

# RAG Evaluation Framework

## Purpose

Use this file to design evaluations that support product and engineering decisions.

Evaluate the layer being changed, then verify end-to-end behavior.

# 1. Start with the decision

Examples:
- Did hybrid retrieval improve answerability?
- Should we change chunking?
- Did the reranker improve ordering enough to justify latency?
- Did a model change improve groundedness?
- Is the new index safe to promote?

Write the decision before choosing metrics.

# 2. Eval dataset

Build from:
- real user queries;
- known failures;
- important business tasks;
- hard negatives;
- ambiguous questions;
- no-answer cases;
- conflicting-source cases;
- multilingual/domain slices;
- security/adversarial cases.

Track expected source/evidence where possible.

# 3. Retrieval evaluation

When relevance labels/qrels exist, use metrics such as:
- recall@k;
- precision@k;
- MRR;
- NDCG;
- hit rate.

Also inspect:
- relevant document missing;
- relevant chunk poorly ranked;
- irrelevant source dominating;
- filter leakage.

Do not judge retrieval only by final answer quality.

# 4. Context evaluation

Measure whether final model context:
- contains required evidence;
- excludes unauthorized evidence;
- avoids excessive duplication;
- preserves provenance;
- fits token budget;
- handles conflicts/recency.

# 5. Generation evaluation

Possible dimensions:
- groundedness;
- factual correctness;
- relevance;
- completeness;
- instruction adherence;
- abstention;
- citation correctness.

Separate answer quality from retrieval quality.

# 6. Citation evaluation

Evaluate:
1. citation presence where needed;
2. claim support;
3. correct source attribution;
4. source coverage for multi-claim answers.

A high citation count is not a quality metric.

# 7. No-answer / abstention

Include cases where the correct behavior is:
- “not enough evidence”;
- “source unavailable”;
- “conflicting sources”;
- “not authorized.”

Measure false answers and unnecessary abstention separately.

# 8. LLM-as-judge

Use model judges for qualitative criteria when needed.

Guardrails:
- explicit rubric;
- representative examples;
- calibration against human labels;
- consistency checks;
- blinded comparison where practical.

Do not treat judge output as ground truth.

# 9. Slices

Report by meaningful slices:
- query type;
- document type;
- language;
- tenant/domain;
- answerable vs unanswerable;
- fresh vs stale;
- short vs complex query.

Avoid hiding a bad slice inside one average.

# 10. Regression gates

Define before launch:
- minimum retrieval quality;
- no security regression;
- groundedness/citation threshold;
- latency budget;
- cost budget;
- no-answer behavior.

Document why a threshold changes.

# 11. Online validation

Use when appropriate:
- shadow;
- canary;
- limited beta;
- feature flag;
- human review.

Offline improvement is not proof of production improvement.

# 12. Continuous evals

Add production failures to the eval set.

Version results by:
- corpus/index;
- embedding model;
- chunker;
- retriever;
- reranker;
- prompt;
- model;
- application version.

This makes regressions diagnosable instead of anecdotal.

SHA-256: aff2b108604334df18a633db9004e4558451a07f8698d8b3efd657882af097a1