A RAG system that feels good in a demo can still fail on citation accuracy and hallucination rate. We gate launches on measurable quality.
Core metrics
We track answer faithfulness, context relevance, citation coverage, and latency. No single metric is enough.
Human review samples
Automate scoring, then spot-check hard cases with domain experts. Edge documents expose retrieval gaps fast.
Failure taxonomies
Classify misses: wrong chunk, missing chunk, bad synthesis, or unsafe content. Fixes differ for each class.
Go-live criteria
Define thresholds per use case. Internal search can tolerate more noise than customer-facing copilots.