AI engineering · 36 of 42

Precision@k, MRR and NDCG

Recall@k cannot see the order

Scroll

Recall@k cannot see the order

Recall@k answers whether the evidence was in the top k at all. It says nothing about where, and position matters enormously when only three documents will fit in the window.

Precision@k asks what fraction of those k were actually relevant — the cost of padding. MRR rewards getting the right thing first and drops sharply as it slides down. NDCG goes further, crediting partially relevant results and discounting by position.

Use more than one. A system with excellent recall and poor MRR is finding the right documents and burying them, which is a reranking problem. Excellent MRR and poor recall is a retrieval problem. One number cannot tell you which you have.

Evaluation
RECALL@K CANNOT SEE THE ORDER evidence at rank 1 1 the answer 2 noise 3 noise 4 noise 5 noise recall@5 = 1 MRR = 1.00 evidence at rank 5 1 noise 2 noise 3 noise 4 noise 5 the answer recall@5 = 1 MRR = 0.20 Same recall, very different experience. MRR rewards being first; NDCG also credits partial matches further down.
The same evidence at rank one and at rank five: identical recall, very different reciprocal rank.