← Files RAG & GenAI CopilotARCHIVED FILE

skills/production-rag-genai-copilot/references/rag_observability_latency_cost.md

2.57 KB · Sep 30, 2026 · 23:18 UTC

↓ Download file

# RAG Observability, Latency, and Cost

## Purpose

Use this file to instrument and optimize production RAG systems without sacrificing correctness blindly.

# 1. Trace the whole request

Capture when relevant:

- request/run ID;
- user/tenant context identifier;
- query;
- rewrite/router output;
- filters;
- retrieval method;
- retrieved document/chunk IDs;
- retrieval scores;
- reranker input/output;
- context IDs/order/token count;
- model/version;
- response/citations;
- latency by stage;
- tokens/usage/cost;
- errors/retries;
- final outcome.

Redact sensitive payloads.

# 2. Latency budget

Break end-to-end latency into:
- rewrite/router;
- retrieval;
- reranking;
- context processing;
- model generation;
- tools/external APIs;
- post-processing.

Optimize the dominant stage first.

# 3. Retrieval latency

Check:
- index/query type;
- filter complexity;
- candidate count;
- hybrid/multi-query fan-out;
- reranking depth;
- network region;
- concurrency.

A faster retriever that destroys recall is not an optimization.

# 4. Model latency

Check:
- model choice;
- input context size;
- output length;
- reasoning level where applicable;
- streaming;
- concurrency/rate limits.

Measure quality per successful task, not only tokens/request.

# 5. Context cost

Reduce context using evidence:
- dedupe;
- better ranking;
- per-source caps;
- parent/child expansion;
- concise metadata;
- dynamic top-k;
- compression only if it preserves evidence.

Do not truncate required evidence merely to lower token cost.

# 6. Reranking trade-off

Measure:
- ranking gain;
- answer-quality gain;
- added p95 latency;
- added cost.

Remove reranking only if first-stage ordering is already sufficient.

# 7. Caching

Possible cache layers:
- parsed documents;
- embeddings;
- retrieval results;
- model responses;
- summaries.

Define:
- cache key;
- tenant/auth scope;
- freshness;
- invalidation;
- privacy.

Never share cached private results across authorization boundaries.

# 8. Indexing cost

Track:
- parsing volume;
- embedding volume;
- re-embedding rate;
- index storage;
- update frequency;
- stale-document cleanup.

Avoid full reindex when a scoped update is sufficient and safe.

# 9. SLOs

Useful production signals:
- availability;
- answer success;
- retrieval recall proxy/eval score;
- citation/groundedness failures;
- p50/p95/p99 latency;
- error/retry rate;
- stale-index lag;
- cost per request/success.

Tie alerts to an action threshold.

# 10. No-action case

If quality, security, latency, and cost are inside targets, avoid speculative optimization.

Define the threshold that would justify change.

SHA-256: c27f494d36dd338da8543e215a463b93e712db66c5d542cb1143013291573a1a