Freeze the boundary
The corpus rows, fixtures, retrieval configuration, and checksums are fixed before comparison.
Measured retrieval evidence
A citation-sensitive RAG system needs more than recall. We test ranking, source anchors, citation focus, abstention, and scope violations against one frozen corpus and fixture set.
Direct answer
Use recall and rank quality to measure retrieval, source-anchor recall and citation precision to measure source alignment, and explicit abstention tests to catch unsupported queries. Track boundary violations separately. A high recall score cannot prove that citations are focused, that unsupported queries abstain, or that the result stayed inside its authorized corpus.
Evaluation framework
The benchmark uses 1,160 indexed document rows and 80 fixed queries. Fifty fixtures calibrate the abstention threshold; 30 remain in validation. The set contains 40 semantic queries, 20 exact references, 10 multi-evidence requests, and 10 unsupported queries. Ranking and citation metrics average the 70 answerable queries; abstention accuracy covers the 10 unsupported queries. The calibration and validation split serves threshold selection, not the denominator of every metric.
The corpus rows, fixtures, retrieval configuration, and checksums are fixed before comparison.
Each fixture identifies relevant chunks and required official source anchors at a cutoff of ten.
Unsupported queries must abstain. Every returned passage must remain inside the enforced source scope.
Metric definitions
These metrics answer different questions. CorpusMesh keeps them separate instead of collapsing retrieval quality into one score.
Production endpoint result
These values come from 80 responses recorded against the deployed private-beta endpoint on 2026-08-21. This is a frozen v1 evaluation, not a measurement of the current corpus or live source freshness.
The score profile is intentionally visible: citation precision is lower than recall. That is useful evidence for improving focus, not a number to hide.
Request beta accessConfiguration selection
Both embedding candidates used the same 1,160 rows and 80-query fixture set. Gemini embedding 2 at 768 dimensions cleared the gate. Voyage 4 at 1024 dimensions did not.
| Configuration | Gate | Recall@10 | nDCG@10 | MRR@10 | Citation precision |
|---|---|---|---|---|---|
| gemini-embedding-2 · 768D | Passed | 96.2% | 82.0% | 78.9% | 53.4% |
| voyage-4 · 1024D | Not selected | 92.6% | 69.9% | 64.7% | 49.1% |
Worked query
The fixed source-trace gate asks how Article 6 classification depends on Annex III. The production trace ranked Article 6 first and Annex III at position 3.
Inspect the request, passages, and EUR-Lex citations →57c28e321dbdReproducibility and limits
The fixture, chunk manifest, runtime report, and deployed commit are recorded so the result can be tied to one evaluated system state. These identifiers support provenance; this page does not provide the complete fixture and corpus artifacts needed for an independent replay.
c1780012a679b35b8d43e32bd8be4149af391be7cd971a66e54d2f348fa65b732d73f00da7bd0cfbf942c26e96dbe27bbcf9161f428dd7985d6d669581e74924052cbcd7621f56e6d37b7addd7c931234af2d0f1422948bc3c29210eb2500c7fYour evaluation worksheet
Use this public benchmark to inspect the method. Define a separate gate for your own sources and application, with an owner and a pass condition for each row.
| Check | Record before evaluation | Evidence to review |
|---|---|---|
| Source scope | Authorities, documents, version, languages, jurisdictions, and exclusions | Every returned passage stays within the approved boundary. |
| Answerable queries | Representative questions and expected article or section anchors | Recall, ranking, and citation focus, with query counts and the chosen cutoff. |
| Unsupported queries | Questions outside the source scope or missing required evidence | Explicit abstention rather than unrelated context. |
| Source changes | Check cadence, candidate review, activation criteria, and recovery owner | A reviewed version, regression results, and a recoverable prior version. |
| Application outcome | Your model, prompts, user workflow, and acceptance criteria | Final-answer quality and usability measured separately from retrieval. |
Citation precision here is 53.4%: high recall still leaves room to improve the focus of retrieved passages. The worked query is one successful trace, not a complete account of every failure. The published benchmark uses ten results; the prepared beta plan caps a request at five, so evaluate your workload at its actual limit.
Review the freshness method and the published beta scope and price before defining your pilot.
Evaluate your workload
Tell us the sources, jurisdictions, query types, and failure modes that matter. We will define the fixture boundary before discussing a production knowledge base.