LLM Observability and RAG Evaluation: Building a Production Feedback Loop
Build a production LLM observability and RAG evaluation loop: trace every request, score outputs, review edge cases, and catch regressions before users do.
Your RAG assistant confidently told a customer that the return window was 90 days. The policy changed to 30 days six weeks ago. The retrieval layer pulled a stale chunk from a deprecated help-center article, the model synthesized it fluently, and the user got a wrong answer delivered with total conviction. Nothing in your APM dashboard fired. Latency was 1.4 seconds, status code 200, zero exceptions. The system was perfectly healthy and completely wrong.
This is the defining failure mode of production LLM systems, and it is why LLM observability without RAG evaluation is just expensive logging — and evaluation without observability is guesswork. You need both, wired together into a loop that catches regressions before your users do.
That loop has four stages: capture → score → act → re-test. Capture every request in enough detail to replay it. Score outputs against automated and human signals. Act on the scores by fixing prompts, retrieval, or data. Re-test to prove the fix holds and nothing else broke.
Three pillars support those stages, and each one alone gives you false confidence:
- Tracing (the capture stage) tells you what happened, but not whether it was good.
- Evals (the score stage, automated half) tell you whether outputs were good, but not why they weren't.
- Human review (the score stage, ground-truth half) gives you labels you can trust, but doesn't scale without the other two.
The act and re-test stages are the operational layer that turns those three pillars into engineering work: a fix, and a test that proves the fix holds. Wire all three pillars together and you get something no single tool provides — the ability to move from "quality dropped 8% on Tuesday" to "the reranker started dropping policy documents after the index rebuild" in minutes.
Why LLM Observability Is a Different Discipline Than APM
Traditional APM was built for deterministic software. A function either returns the right value or throws. A service is either up or down. Uptime, p99 latency, and error rate are sufficient because the contract is binary and stable.
LLM systems break every one of those assumptions.
Non-determinism is the default, not an edge case. The same prompt at temperature 0.7 can produce materially different answers across calls. Sampling, provider-side batching, hardware variance, and floating-point reduction order all introduce drift you cannot reproduce by replaying a request ID.
Here is the uncomfortable corollary: setting temperature to 0 does not make you deterministic. On modern serving stacks, mixture-of-experts routing, continuous batching, and non-associative floating-point accumulation mean the same input can still yield different tokens on different runs. Treat temperature 0 as "lower variance," not "reproducible." If you need bit-exact replay, you must cache and replay the actual model output, not re-invoke the model.
Prompt sensitivity compounds this. Moving a single instruction from the system prompt to the user turn, or reordering few-shot examples, can shift accuracy by double digits. There is no compiler to catch it and no unit test that covers it unless you wrote one deliberately.
Provider-side model drift is the quietest failure. When a vendor silently updates a model behind an alias like gpt-4o or claude-sonnet, your quality distribution moves and your code does not change at all. This is a well-documented class of incident: teams have shipped multi-week regressions with no code deploy and no alert. Two mitigations matter. First, pin explicit model versions in production rather than floating on an alias, and treat an alias bump as a change that requires a full eval run. Second, detect drift by measurement, not by changelog: run a frozen golden set against the pinned model on a schedule and alert on distribution shift in the score (a Kolmogorov–Smirnov test or a population stability index on the score histogram is enough to start).
The consequence: quality is a distribution, not a boolean. You cannot assert correctness with a status code. You need scored samples, tracked over time, segmented by prompt version, model version, and retrieval index version.
And you must instrument the distribution, not just its mean. Define a per-request quality score, then track its p50 and p10 over time. Regressions hide in the tail: a change can leave the median untouched while the worst 10% of answers get dramatically worse. Alert on the p10, not the average.
One statistical guardrail before you page anyone: with low daily volume, an "8% drop" may be noise. Set a minimum sample size, or use a sequential test, before an alert becomes a page. A feedback loop that cries wolf gets muted — and a muted loop is worse than none.
Finally, because every token costs money and every token adds latency, cost and latency are first-class quality metrics, not infrastructure footnotes. A retrieval strategy that doubles recall but triples context length is not obviously an improvement — you need the numbers side by side.
The Three Pillars of a Production AI Feedback Loop
Tracing: Reconstructing the Full Request Path
A RAG request is not one call. It is a span tree: user query → query rewrite → embedding → vector search → metadata filter → reranker → prompt assembly → model call → optional tool calls → response. If you only log the final prompt and completion, you have thrown away the evidence you need to debug.
Capture per span: inputs and outputs, model name and version, prompt template version, token counts (input, output, cached), latency, and — for retrieval spans — the retrieved document IDs with their similarity scores and ranks. Attach a correlation ID at the edge and propagate it through every hop.
That correlation ID is the difference between a debuggable system and an anecdote. When a user reports a bad answer three days later, you can pull the exact span tree, see that the reranker ranked the correct chunk 14th, and reproduce the failure deterministically by freezing the retrieval results and replaying the generation step.
Two practical notes. First, redact PII at the logging boundary, not at the storage boundary — you want a single, auditable redaction point. Second, store full prompts and completions for a sampled subset rather than everything; full-fidelity traces are the most expensive data you will own. Keep the metadata and scores for every request; keep the raw text for a sample.
Evals: Turning Outputs Into Pass/Fail Signals
Evals come in two families, and mature teams run both.
Reference-based evals compare outputs to golden answers. You curate a dataset of 100–500 representative queries with verified correct responses or verified correct source documents, then score exact match, semantic similarity, or rubric compliance. These are precise and cheap to run repeatedly, but they only cover what you thought to include.
Reference-free evals score outputs without a gold answer. LLM-as-judge grades faithfulness or relevance against a rubric; heuristics check format, citation presence, and refusal behavior; embedding similarity measures drift from a baseline. These scale to production traffic, which is exactly why they are dangerous: a broken judge will confidently mislabel thousands of requests.
Calibrate the judge before you trust it. LLM-as-judge is not a primitive you install; it is a model you must validate. Three failure modes dominate:
- Position bias — judges favor whichever candidate appears first. Always run both orderings and average, or randomize.
- Verbosity and self-preference bias — judges favor longer answers and outputs from their own model family. Use a judge from a different family than the system under test where possible, and control for length.
- Rubric drift — an unversioned rubric silently changes meaning over time. Version the rubric like code.
The non-negotiable step: build a small human-labeled calibration set (a few hundred examples) and measure judge agreement against it — report Cohen's κ or a correlation, not a vibe. If your judge doesn't agree with humans on the calibration set, its scores on production traffic are noise. Re-calibrate whenever you change the judge model or rubric.
Human Review: The Ground Truth That Scales Through Sampling
Human review is where you get labels you can actually trust — and where most teams either over-invest (reviewing everything) or under-invest (reviewing nothing). The answer is stratified sampling, not exhaustive review.
Route to humans in three cases:
- Low-confidence automated scores. When the judge score sits in an ambiguous band, or heuristics disagree with the judge, escalate.
- High-stakes or high-visibility requests. Flag by user tier, topic (billing, legal, safety), or any request that triggered a refusal or a complaint.
- A random control sample. Always review a small random slice of passing requests. This is the only way to catch the failures your automated scores are blind to — the false negatives that make your dashboard look green while users are unhappy.
Give reviewers a structured rubric and a span-level view, not just the final answer: they need to see the retrieved chunks to distinguish a retrieval failure from a generation failure. Log every human label back into the same store as the automated scores so the two can be compared — and periodically check whether your automated scores are drifting away from human judgment.
Closing the Loop: From Failure to Regression Test
The loop only pays off when a failure becomes a permanent test. The workflow:
- Capture the failure as a span tree with a correlation ID.
- Diagnose the layer — retrieval, reranking, prompt, or model. The span tree tells you which.
- Fix it — update the index, adjust the reranker, revise the prompt, or pin a model version.
- Freeze the case into the eval suite. Add the query, the expected source documents, and the acceptance criterion to your golden dataset. This is the step teams skip, and it is the entire point.
- Gate deploys on the suite. Every prompt change, retrieval change, and model bump runs the full eval suite in CI. No eval pass, no deploy.
Two guardrails keep the loop honest. First, treat the eval suite as code: version it, review changes to it, and be suspicious of anyone who loosens a threshold to make a build pass. Second, watch cost and latency in the same gate — a prompt change that improves quality by 2% but doubles token spend is a trade-off decision, not an automatic win.
What to Instrument First
If you are starting from nothing, do it in this order:
- Correlation IDs and span trees for every request. Without this, nothing else is debuggable.
- Per-request quality scores, tracked as p50 and p10, segmented by prompt, model, and index version.
- A golden dataset of 100–500 cases with verified sources.
- A calibrated judge with a human-labeled agreement set.
- A CI gate that runs the eval suite on every deploy.
The failure mode that opened this article — the fluent, confident, wrong answer — is invisible to APM by design. It is only visible to a system that captures the full request path, scores the output, escalates the ambiguous cases to a human, and turns every caught failure into a test that runs forever. Build the loop, and the next stale chunk gets caught in CI instead of in front of your customer.