Your reranker bake-off is a tie
As rerankers converge, nDCG on coarse human labels stops telling them apart. Rubric-Calibrated Preferences lets an LLM compare, a rubric anchor, and IRT put every query on one scale.
Every team choosing a reranker runs the same bake-off. Take a benchmark, swap in Cohere, Voyage, Jina, a Qwen3 reranker, or the in-house fine-tune, and compare nDCG@10. The numbers land within a point or two of each other. Most pairs are “not significant.” Someone picks the cheapest one and calls it a decision.
That tie is often not a fact about the rerankers. It is a fact about the labels. Schmidt, Crisostomi, Lassance, and Reimers (Cohere, September 2026) offer a fix (arXiv:2609.35739). Rubric-Calibrated Preferences (RCP) uses an LLM to compare documents, a rubric of yes/no criteria to anchor them, and Item Response Theory to put every query’s judgments on one shared scale. The resulting metric, RCP-nDCG, separates rerankers that nDCG on human labels cannot — and, on the paper’s human checks, separates them in the right direction.
Why nDCG runs out of resolution
nDCG scores a ranking against relevance labels (“qrels”): relevant or not, sometimes graded 0–3. Those labels are costly, sparse, and noisy, and any document nobody judged counts as irrelevant. As rerankers converge, per-query nDCG fails in three ways: it saturates when every reranker puts the one labeled document on top, floors when none does, and compresses when most pairs tie. Across the 14 rerankers the authors evaluate, one of those failures hits 45.6% of NanoBEIR queries and 30.7% of BRIGHT queries. On one FiQA query, 13 of 14 rerankers tie at 1.0 although their lists differ.
LLM judges are the obvious source of denser labels, but each standard mode breaks something. Relative judgments — “A beats B for this query” — tell close documents apart, but their scores mean nothing across queries: a Bradley–Terry score is fixed only up to a constant, and a judge can be more decisive on one query than another. Absolute grades share a scale but are too coarse; many documents land on the same grade. You get resolution or comparability. nDCG needs both.
What RCP does
RCP makes two passes over a benchmark’s candidate pool, then merges them.
Stage A is a listwise tournament. The judge scores windows of ten documents, the pairwise preferences inside each window are kept, and a Bradley–Terry model turns them into one fine-grained order per query.
Stage B is a rubric. The judge checks each document against five yes/no criteria of increasing stringency: topical relevance, useful information, entity/detail match, a direct answer, and thorough treatment. Identical criteria for every query make an absolute standard.
The merge borrows from psychometrics: IRT puts test-takers on one scale using shared questions. Here the documents take the test and the criteria are the questions. A two-parameter IRT model learns each criterion’s difficulty and discrimination once, across all queries, and fits a per-query scale and offset that map that query’s tournament scores onto the shared axis. Calibration provably keeps the order inside a query; it only fixes how documents from different queries compare. The calibrated score becomes a continuous gain between 0 and 1, and RCP-nDCG is ordinary nDCG with that gain in place of qrels.
The cost structure suits evaluation. The pool is judged once; RCP-nDCG then scores any reranker that reorders it without further judge calls. The priced TREC-DL run, with the open-weight Qwen3.6-27B judge at public API rates, came to $3.72 per query.
What the evidence says
Forty-six paid external annotators graded the union of two rerankers’ top-five lists, blind to which list each document came from: 7,080 grades on 2,268 documents.
- Calibration does real work. The Spearman correlation between a query’s mean score and its annotators’ mean grade rises from 0.538 with raw tournament scores to 0.795 after calibration.
- The labels beat the qrels. Given one document annotators rate useful and one they don’t, the RCP gain ranks the useful one higher with AUC 0.910; the benchmark’s qrels manage 0.651, largely because they give about half such pairs the same label.
- Rerankers are compared better. In 185 contests where exactly one metric agrees with the annotators’ verdict, that metric is RCP-nDCG 72.4% of the time. Chance is about 53%, not 50%, because qrel-nDCG ties more often.
- It does not fight NIST. On TREC-DL, RCP-nDCG favors NIST’s winner on every reranker pair that NIST’s grades separate significantly, with either of two judges. That is 51 pairs — 14 of 91 in 2019, 37 of 91 in 2020 — and RCP-nDCG separates many more.
- It breaks ties. On NanoBEIR, the share of the 91 reranker pairs a paired t-test separates rises from 34.1% to 63.5% — the abstract’s “1.9×.” Most of that added separation comes from the rubric; the tournament orders documents the rubric alone would tie.
The leaderboard moves. On NanoBEIR the ZeroEntropy rerankers sit last under qrel-nDCG@10; under RCP-nDCG@10, zerank-2 and zerank-1 rank first and second. Those models train on LLM preferences, so that is exactly the result to distrust. The authors check it: in 66 human contests where the two metrics split on a zerank pair, annotators side with zerank 78.8% of the time.
Why this matters if you buy or build rerankers
If you are choosing among vendor rerankers or your own fine-tunes, you are probably deciding inside nDCG’s noise floor. RCP gives a defensible way to use an LLM judge for that call without the standard objection that its scores don’t compare across queries. The pedigree helps — Reimers wrote Sentence-Transformers; Lassance worked on SPLADE — but the disclosure matters more. Cohere funded the work and sells two of the 14 rerankers tested; Rerank 4 Pro lands fifth under RCP-nDCG@10 on NanoBEIR.
This is the sequel to Calibrate to the assessor, which argued that relevance is a calibration problem and a prompted LLM is no drop-in replacement for a human pool. RCP accepts that premise rather than dodging it. Raw tournament scores are miscalibrated, and the paper says so: expected calibration error of 0.371 against the judge’s own rubric answers, 0.017 after the IRT fit. The recommendation is conservative: label the pool once and report RCP-nDCG next to qrel-nDCG, not instead of it.
What not to claim
The rubric targets general-purpose retrieval; other domains may need other criteria. An ablation over fifteen rubric variants on TREC-DL reversed no significantly separated pair unless the rubric ignored the query or bunched every criterion at one end — robust to wording, not a license to skip design.
Gains are pointwise. They ignore redundancy and complementarity, so they say nothing about whether a RAG context set covers a multi-hop question. They measure perceived relevance, not factual correctness.
The human reference is noisy too. Two annotators grade a document within one step of each other 77.8% of the time, and their questions resemble the rubric by design. Validation covers NanoBEIR, BRIGHT, ViDoRe v3, and TREC-DL — not your corpus.
Scores are comparable only within one calibration fit. New documents need new judgments and a refit, which can shift earlier scores, so pin the judge and rubric and re-baseline when either changes. A second judge reproduces the leaderboard closely (mean Spearman 0.959 on NanoBEIR), but LLM judges may share the rerankers’ preferences — never train on the labels you evaluate with. And treat RCP-nDCG@5 margins under 0.02 as ties; annotators agree with them only at chance.
Build implication
Stop declaring a reranker winner from nDCG deltas on sparse binary labels. Build the candidate pool once. Have a pinned, open-weight judge run a listwise tournament and a fixed yes/no rubric over it, fit the IRT calibration, and score every candidate against the calibrated gains. Keep qrel-nDCG alongside: treat agreement as confidence and disagreement as the place to spend human review.
Your bake-off is a tie because your labels are coarse. Compare with the LLM, anchor with the rubric, calibrate across queries — then decide.
Paper: Schmidt, Crisostomi, Lassance, and Reimers. Code and data: cohere-ai/rcp-ndcg.