Every technique in this series makes your knowledge layer better, or claims to. Evaluation is how you tell the difference, and it’s the discipline most teams skip and most regret skipping.
The problem
A team has shipped the whole stack from this series. Hybrid search, a reranker, a knowledge graph, an agentic loop, a memory system. Each one was added because it “felt better” in a demo. Then a stakeholder asks a fair question: the reranker doubled our latency: what did it buy us? And the room goes quiet, because nobody can put a number on it. The system might be genuinely good. It might also be carrying three expensive components that move no metric at all, and a couple of regressions that nobody noticed because there was nothing watching.
That silence is the failure this article is about. You cannot manage a knowledge layer you don’t measure, and most teams measure it on vibes: a few hand-tried questions before a release, a thumbs-up from whoever built it. Evaluation is the part of the knowledge layer (the part of an agent system that decides reliability in production, and an engineering discipline in its own right) that turns “feels better” into “recall@5 went from 0.71 to 0.86 and faithfulness held.” Without it, every other article in this series is a guess you’re paying for.
How evaluation works
The spine: measure the layer at two altitudes (offline against a fixed golden set before you ship, and online against live traffic after) and measure retrieval and generation as separate things, because they fail separately
Figure 1: Evaluation as a loop. A golden set drives offline retrieval and generation metrics that gate releases; live instrumentation catches what the golden set didn’t, and feeds new cases back into it.
Start with retrieval versus generation, because conflating them hides the bug. Retrieval metrics ask: did the right context come back? That’s recall@k, and context precision / context recall (did the retrieved chunks contain what was needed, and was the needed thing retrieved). Generation metrics ask: given that context, was the answer good? That means faithfulness (is the answer grounded in the retrieved context or hallucinated) and answer relevance (does it address the question). RAGAS (Shahul Es and co-authors, RAGAS: Automated Evaluation of Retrieval Augmented Generation, [arXiv 2309.15217](https://arxiv.org/abs/2309.15217), EACL 2024) is the reference toolkit that operationalised exactly this split, and it’s the right starting vocabulary even if you build your own harness. The reason the split matters: a faithfulness drop with stable retrieval is a generation problem; a recall drop with stable faithfulness is a retrieval problem; and if you only measure the final answer you can’t tell which knob to turn.
Then layer in the workload-specific benchmarks, because a knowledge layer with memory or long histories needs eval the RAG metrics don’t cover. LongMemEval (Di Wu and co-authors, [arXiv 2410.10813](https://arxiv.org/abs/2410.10813), ICLR 2025) tests long-term memory across sessions (extraction, temporal reasoning, knowledge updates) and is the benchmark from Article 6 that put numbers on memory degradation. MemBench (Haoran Tan and co-authors, [arXiv 2506.21605](https://arxiv.org/abs/2506.21605), Findings of ACL 2025) pushes on memory effectiveness, efficiency, and capacity together. Pick the benchmark that matches your workload; don’t measure a memory agent with a single-shot RAG metric and call it covered.
In practice
Take a concrete, public example: [srb-legal-rag](https://github.com/RatkoNikolic/srb-legal-rag), the open benchmark from the [series special](https://www.ratkonikolic.com/p/serbian-legal-ai-a-rag-architecture) (August 2026). Five retrieval architectures answer the same 60 Serbian legal questions over the same knowledge base, with the generator, prompts and chunking held constant so retrieval is the only thing that changes. Every question carries a known-correct answer and the exact statute articles it must cite, and the set is tiered by difficulty: 20 easy, 20 medium, 20 hard.
Read the table without the statistics and you’d conclude the graph beat advanced RAG by four points. Read it with them and you conclude nothing of the sort: the two arms disagree on only six questions, split 4 to 2, and a paired McNemar test gives p = 0.69. That’s noise, not an effect. The agentic arm’s +0.13 over advanced looks like a win at p = 0.039, but across six pairwise comparisons it doesn’t survive multiple-comparison correction. What does survive is the jump from naive to every engineered arm (p ≤ 0.0003), and agentic’s quality profile: lowest hallucination, best provenance, zero wrongful refusals, at about 2.3 times the cost and nearly four times the p95 latency of advanced. That’s the real lesson of the table: the numbers that look like a result and the numbers that are a result are different sets, and only paired significance testing tells you which is which.
Two more things the harness taught that apply to any team. The judge gets graded too. Answers are scored by an LLM judge from one model family and re-graded by a juror from another; their agreement came in at Krippendorff’s α = 0.764, below the usual 0.80 bar. Most of the disagreement turned out to sit on one boundary, “correct” versus “partially correct” (ordinal α = 0.842), and re-deriving the ranking from the juror’s labels changed nothing. That’s a publishable caveat instead of a hidden one. Small cells lie. The temporal tier, the one built to test “which version of the law was in force on this date,” has three questions and only one of them actually separates the arms: an effective n of 1. It demonstrates a mechanism; it doesn’t measure one. And the easy tier is already saturated for every engineered arm (19 or 20 out of 20), so all the discrimination lives in medium and hard. Provenance and temporal fidelity are first-class columns here for a reason: in a high-stakes domain a confidently wrong, uncited answer is the failure that matters, and a single accuracy number would hide it completely.
For the online half, instrumentation replaces the golden set’s certainty with live coverage: log retrieval traces and answers, sample them, watch for drift (recall quietly sliding as the corpus grows), and capture user signals (thumbs, escalations, re-asks) as the cheapest labels you’ll ever get. The tooling here has matured: LangSmith, Langfuse, and Braintrust are the current platforms teams use to trace, score, and regression-test agent outputs in production, and DeepEval alongside RAGAS covers the open-source CI side. You don’t need all of them; you need one, wired in before launch rather than after the first incident.
Where it shines
Two places evaluation stops being hygiene and becomes leverage.
The first is change decisions: every “should we add X” in this series. A golden set turns the reranker question, the graph question, the agentic question into measured deltas instead of debates. The team that can say “the reranker moved recall@5 +0.15 for +400ms p95” makes better architecture calls than the team arguing from demos, and it makes them faster, because the evidence is already on the table.
The second is catching silent regressions: the failures that don’t throw errors. A corpus update that quietly tanks recall for one query class, a model upgrade that subtly changes faithfulness, drift that accumulates over weeks. None of these page anyone; they just degrade the product until a customer complains. Continuous evaluation is the only thing that sees them coming, and it’s the difference between finding a regression in CI and finding it in a churn report.
Where it breaks
Evaluation has its own failure modes, and pretending it doesn’t is how teams get falsely confident.
Golden sets rot. A fixed question set made once and never touched stops reflecting real traffic within months: new question types appear, the corpus shifts, and you’re optimizing against a museum. The golden set is a maintained asset, not a one-time artifact, and the maintenance is the part teams skip.
Benchmarks saturate and stop discriminating. When everything scores 0.95 on your benchmark, the benchmark has stopped telling you anything. That’s Goodhart’s law, live: once a measure becomes the target, it stops being a good measure. High benchmark numbers can coexist with mediocre real-world reliability, which is precisely the benchmark-to-production cliff the knowledge-editing article ran into. A score that only goes up is often a sign the test got easy, not that the system got good.
LLM-as-judge has biases of its own. Using a model to grade outputs is now standard practice (it’s how you score faithfulness at scale without armies of annotators), but the judge is not neutral. The canonical study (Lianmin Zheng and co-authors, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, [arXiv 2306.05685](https://arxiv.org/abs/2306.05685), NeurIPS 2023) documented position bias, verbosity bias, and self-preference, and later work (Lin Shi and co-authors, [arXiv 2406.07791](https://arxiv.org/abs/2406.07791), AACL-IJCNLP 2025) systematised just how strong the position bias is across judges. The judge is useful and it needs its own evaluation: calibrate it against human labels on a sample before you trust it to grade thousands.
When not to use it
The honest answer is almost never, but the level of investment should match the stakes. A throwaway internal prototype for five users doesn’t need a maintained golden set, an LLM-judge harness, and online instrumentation; a few tried questions before each change is proportionate, and building a full eval rig for it is procrastination dressed as rigour. The moment the system touches customers, money, or compliance, that flips hard: the cost of a silent regression now dwarfs the cost of the harness, and “we’ll add eval later” becomes the line teams quote ruefully after the incident. Match the eval to the blast radius.
Cost / latency / setup effort:
Combinations and hybrids
Evaluation isn’t a technique that composes with the others: it’s the substrate that lets you choose among them rationally. Every prior article ends with a “measure it” checklist item, and this is the article those items point to: the reranker delta, the graph’s provenance gain, the agentic loop’s accuracy-vs-latency trade, the memory system’s recall. All of it is a number a golden set produces or a debate that never ends. The strongest setup pairs offline gating (catch regressions before release) with online instrumentation (catch what the golden set didn’t), feeding live failures back into the golden set so it tracks reality instead of rotting. That loop is what makes the whole series operational rather than aspirational, and it’s the natural lead-in to the final article, which uses exactly this kind of measured comparison to choose a knowledge layer for a given problem.
Production checklist
Build a golden set early (50–200 real questions with known-good answers and source chunks) and treat it as a maintained asset, not a one-off.
Measure retrieval and generation separately so a bad answer points at the right subsystem (recall vs faithfulness).
Gate releases on the golden set: no change to retrieval, prompt, or model ships without the before/after delta.
Instrument production: log traces, sample answers, capture thumbs/escalations/re-asks as labels.
Watch for drift: track recall and faithfulness over time, not just at release; the corpus moves under you.
Calibrate your LLM judge against human labels on a sample before trusting it, and re-calibrate on judge-model upgrades.
Add domain metrics that matter: for regulated work, provenance and temporal validity are first-class, not afterthoughts to an accuracy number.
Refresh the golden set on a schedule with real failures from production, so it tracks traffic instead of aging into a museum.
The take
You cannot improve a knowledge layer you don’t measure, and “it felt better in the demo” is not measurement. Build a golden set, separate retrieval from generation, gate every change on the numbers, instrument production for the regressions the golden set misses, and treat your LLM judge as another component that needs evaluating. It’s the least glamorous discipline in the series and the one that makes all the others honest, because the difference between a knowledge layer that’s reliable and one that merely feels reliable is entirely in whether you can prove it.
I take on a small number of advisory engagements each year for teams hitting exactly these problems. Reach out if that’s you.
What to read next
This series:
Article 1 — Your agent’s problem isn’t the model, it’s the knowledge layer
Article 2 — Naive RAG: the baseline you’ll always benchmark against
Article 3 — Advanced RAG: what you actually run in production
Article 4 — GraphRAG and friends: when entities and relationships beat similarity
Article 5 — Agentic RAG: when retrieval becomes a decision, not a pipeline
Article 6 — Vector vs graph vs episodic: a tour of agent memory systems
Article 7 — Context engineering: the discipline that replaces prompt engineering
Article 8 — Knowledge editing and why it isn’t the answer for enterprise updates
Article 9 — Evaluation: how do you know any of this is working?
Article 10 — Choosing your knowledge layer: a practical enterprise selection guide (coming soon)
External:
Es et al. (EACL 2024), RAGAS: Automated Evaluation of Retrieval Augmented Generation. https://arxiv.org/abs/2309.15217. The retrieval-vs-generation metric split, operationalised.
Wu et al. (ICLR 2025), LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. https://arxiv.org/abs/2410.10813. The benchmark for memory systems.
Zheng et al. (NeurIPS 2023), Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. https://arxiv.org/abs/2306.05685. Why the model-judge is useful and biased, and needs its own evaluation.
Tan et al. (Findings of ACL 2025), MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents. https://arxiv.org/abs/2506.21605. A recent, broader memory benchmark.




