Retrieval Pipeline
retrieved (not in final top-10)
final top-10 · relevant
final top-10 · irrelevant
pool only (not retrieved)
Recall@10
--
nDCG@10
--
Latency (ms)
--
Relevant found
--
How it works

A retrieval pipeline first pulls a broad candidate pool by embedding similarity (angle from the query = distance in embedding space), then spends a limited reranking budget re-scoring the most promising candidates with a more expensive, more accurate model. Radius from the center encodes each document's current score — closer means a higher score at the current pipeline stage.

Recall@k = |relevant ∩ top-k| / min(k, |relevant|)
DCG@k = Σ (2^rel_i − 1) / log2(i + 1)
nDCG@k = DCG@k / IDCG@k
  • Retrieval pool (K0) — how many candidates the first-stage retriever pulls; too small and relevant documents never even reach the reranker.
  • Rerank budget (N) — how many of the pooled candidates get the expensive cross-encoder pass; the rest keep their cheap similarity score.
  • Hard negative density — irrelevant documents that sit close to the query in embedding space, inflating the first-stage score and crowding out true matches.
  • Reranker strength — how well the reranker separates true relevance from surface similarity; low strength leaves hard negatives in the final list.

Raising the rerank budget or strength improves Recall@10 and nDCG@10 but drives latency up — the same latency-quality trade-off production RAG systems have to tune.