Measured retrieval evidence

How we evaluated retrieval quality on 80 EU AI Act queries.

A citation-sensitive RAG system needs more than recall. We test ranking, source anchors, citation focus, abstention, and scope violations against one frozen corpus and fixture set.

Queries80 fixed fixtures
Corpus rows1,160
Evaluation date2026-08-21
Public statusEngineering evidence

Direct answer

Which RAG evaluation metrics matter?

Use recall and rank quality to measure retrieval, source-anchor recall and citation precision to measure source alignment, and explicit abstention tests to catch unsupported queries. Track boundary violations separately. A high recall score cannot prove that citations are focused, that unsupported queries abstain, or that the result stayed inside its authorized corpus.

Evaluation framework

One corpus. One sealed fixture set. Several failure modes.

The benchmark uses 1,160 indexed document rows and 80 fixed queries. Fifty fixtures calibrate the abstention threshold; 30 remain in validation. The set contains 40 semantic queries, 20 exact references, 10 multi-evidence requests, and 10 unsupported queries. Ranking and citation metrics average the 70 answerable queries; abstention accuracy covers the 10 unsupported queries. The calibration and validation split serves threshold selection, not the denominator of every metric.

01

Freeze the boundary

The corpus rows, fixtures, retrieval configuration, and checksums are fixed before comparison.

02

Measure ranked evidence

Each fixture identifies relevant chunks and required official source anchors at a cutoff of ten.

03

Test refusal behavior

Unsupported queries must abstain. Every returned passage must remain inside the enforced source scope.

Metric definitions

Recall alone is not enough.

These metrics answer different questions. CorpusMesh keeps them separate instead of collapsing retrieval quality into one score.

Recall@10
How much of the required evidence appeared in the first ten retrieved passages.
nDCG@10
Whether relevant evidence appeared near the top, with stronger credit for better ordering.
MRR@10
How early the first relevant passage appeared for each answerable query.
Source-anchor recall@10
Whether the expected official article or annex anchor appeared in the result set.
Citation precision@10
What share of retrieved passages matched a required citation anchor. It measures focus, not factual correctness.
Abstention accuracy
Whether unsupported queries returned an explicit abstention instead of unrelated context.

Production endpoint result

The selected configuration passed the sealed gate.

These values come from 80 responses recorded against the deployed private-beta endpoint on 2026-08-21. This is a frozen v1 evaluation, not a measurement of the current corpus or live source freshness.

96.2%Recall@10
82.0%nDCG@10
78.9%MRR@10
98.3%source-anchor recall
53.4%citation precision
100%abstention accuracy
0boundary violations

The score profile is intentionally visible: citation precision is lower than recall. That is useful evidence for improving focus, not a number to hide.

Request beta access

Configuration selection

The alternative did not pass.

Both embedding candidates used the same 1,160 rows and 80-query fixture set. Gemini embedding 2 at 768 dimensions cleared the gate. Voyage 4 at 1024 dimensions did not.

ConfigurationGateRecall@10nDCG@10MRR@10Citation precision
gemini-embedding-2 · 768DPassed96.2%82.0%78.9%53.4%
voyage-4 · 1024DNot selected92.6%69.9%64.7%49.1%

Worked query

Article 6 and Annex III must appear together.

The fixed source-trace gate asks how Article 6 classification depends on Annex III. The production trace ranked Article 6 first and Annex III at position 3.

Inspect the request, passages, and EUR-Lex citations →
Article 6 rank
1
Annex III rank
3
Fixture split
50 calibration / 30 validation
Runtime commit
57c28e321dbd

Reproducibility and limits

What the benchmark proves, and what it does not.

The fixture, chunk manifest, runtime report, and deployed commit are recorded so the result can be tied to one evaluated system state. These identifiers support provenance; this page does not provide the complete fixture and corpus artifacts needed for an independent replay.

Fixture set SHA-256
c1780012a679b35b8d43e32bd8be4149af391be7cd971a66e54d2f348fa65b73
Chunk manifest SHA-256
2d73f00da7bd0cfbf942c26e96dbe27bbcf9161f428dd7985d6d669581e74924
Runtime report SHA-256
052cbcd7621f56e6d37b7addd7c931234af2d0f1422948bc3c29210eb2500c7f

Known limits

  • This benchmark covers one English EU AI Act corpus and its fixed fixture set.
  • Citation precision measures required-anchor focus. It does not prove legal correctness.
  • The results do not establish performance for another domain, language, corpus version, or customer workload.
  • The public evidence does not expose a public retrieval endpoint or legal advice service.

Your evaluation worksheet

Agree the acceptance criteria before the pilot.

Use this public benchmark to inspect the method. Define a separate gate for your own sources and application, with an owner and a pass condition for each row.

CheckRecord before evaluationEvidence to review
Source scopeAuthorities, documents, version, languages, jurisdictions, and exclusionsEvery returned passage stays within the approved boundary.
Answerable queriesRepresentative questions and expected article or section anchorsRecall, ranking, and citation focus, with query counts and the chosen cutoff.
Unsupported queriesQuestions outside the source scope or missing required evidenceExplicit abstention rather than unrelated context.
Source changesCheck cadence, candidate review, activation criteria, and recovery ownerA reviewed version, regression results, and a recoverable prior version.
Application outcomeYour model, prompts, user workflow, and acceptance criteriaFinal-answer quality and usability measured separately from retrieval.

Citation precision here is 53.4%: high recall still leaves room to improve the focus of retrieved passages. The worked query is one successful trace, not a complete account of every failure. The published benchmark uses ten results; the prepared beta plan caps a request at five, so evaluate your workload at its actual limit.

Review the freshness method and the published beta scope and price before defining your pilot.

Evaluate your workload

Test the same retrieval contract against the evidence your product needs.

Tell us the sources, jurisdictions, query types, and failure modes that matter. We will define the fixture boundary before discussing a production knowledge base.

Request beta accessReview the retrieval API