Calibrate to the assessor
SIGIR’s Best Student Paper this year is not another LLM judge prompt. It is a 26M LoRA that learned one assessor, on one topic, and beat the prompted frontier model.
If you build search or RAG systems, you eventually hit the same evaluation hole. You have a test collection. You have human judgments for the documents the original systems retrieved. A new retriever surfaces documents nobody labeled. What do you do with those?
The fashionable answer is to prompt an LLM. Fill the holes. Keep the leaderboard moving. SIGIR’s Best Student Paper this year is not another LLM judge prompt. It is a 26-million-parameter LoRA that learned one assessor, on one topic, and beat the prompted frontier model.
Relevance is a calibration problem, not a reasoning problem. Prompted LLM-as-judge is the wrong default for filling unjudged documents in a reusable IR collection. That sounds like an IR-community quarrel. It is also the same failure mode agent teams hit when they treat a general model’s thumbs-up as ground truth.
The hole in the pool, in plain English
Classic IR evaluation — the Cranfield paradigm behind TREC and a lot of what followed — works like this. Assessors judge relevance for a topic. They cannot judge every document in a large corpus, so you pool: take the top results from many systems, merge them, and judge that set. Documents outside the pool are treated as non-relevant for scoring. That approximation is fine when every competitive system’s good documents made it into the pool. It ages badly.
Cranfield pools go stale. A new system retrieves documents the original assessors never saw. Treating those as non-relevant is unfair to the new system and slowly kills the collection. You either re-judge — expensive, slow, human — or you invent labels for the holes. The fashionable fix is to prompt an LLM to fill them.
That is circular: the same model family can rank and judge. It is prompt-fragile. It is biased toward its own generations. Dietz et al. have been saying this for a while. Gienapp, Potthast, Yates, Scells, and Yang (Kassel, Johns Hopkins, Tübingen; SIGIR 2026, Melbourne) built the alternative.
They fine-tune monoT5 with a per-topic LoRA on one assessor’s pool. They deliberately overfit to that assessor’s notion of relevance for that topic. About 128 labels per topic is enough to beat treating unjudged documents as non-relevant. The adapters are 26 million parameters, against prompted judges up to 229 billion. They are deterministic, they run on a consumer GPU, and they are shareable artifacts. Human judgments stay gold. The adapter transfers the assessor. It does not replace them with a general LLM.
If you come from Solr or OpenSearch production work, the mental model is closer to calibrating a judge to a known assessor’s past decisions than to asking a smart intern what “relevant” means in the abstract. The intern may be fluent. The assessor had a standard.
The body, not the abstract
System rankings from the adapters: Spearman ρ ≥ 0.84 at k=10, and ρ ≥ 0.96 for nDCG, precision, and recall at k=50, against human ground truth. Krippendorff’s α is 0.81–0.88 for adapters against 0.32–0.58 for prompted LLMs. The LLMs preserve system order somewhat, but they systematically over-label relevant — a lenient threshold, applied evenly, so the ranking looks fine while the labels do not. Reasoning-capable models did not win. The authors take that as evidence it is calibration, not reasoning.
That last point is the one I want AI product teams to hear. A “better reasoner” is not automatically a better relevance labeler for an existing pool. Matching one human’s judgments is a different job from writing a fluent explanation of why a document might matter. The paper’s numbers live in agreement and ranking correlation, not in a vibes-based win for chain-of-thought.
Guardrails, also in the body. Do not put the adapter in the ranker. That is training on the test set. Do not transfer it across topics. Do not distill its labels into a retrieval model. This is evaluation infrastructure, not a system component. The LoRA is how you extend a human pool fairly. It is not a free relevance model for production retrieval.
The attack is coming from IR and from stats at once
Gienapp is not a lone dissent. Otero and Parapar (also SIGIR 2026) argue hybrid pooling should keep humans on the shallow pool and let the model label deeper — humans where the ranking decisions are sharpest, models where the pool would otherwise go dark. Lee, Zeng, Jeong, Sohn, and Lee show that naive judge scores are biased; they apply a Rogan–Gladen correction and confidence intervals, because a raw “percent correct” from an imperfect judge is not a measurement. Fiedler (May 2026) shows that sharing calibration across models can sign-reverse a comparison — MMLU-Pro biology, confidently the wrong way. The IR community and the stats community are saying the same thing: “just use GPT as judge” is not an eval.
Then get back to Gienapp. The Best Student Paper is the IR community’s 2026 answer: calibrate to the assessor you have, on the topic you have, with a small adapter you can share and audit. Do not pretend a frontier prompt is a drop-in replacement for a human pool.
ACE made the same point on the agent side: without reliable feedback, the playbook degrades. Spurious lessons stick. A judge without a human pool is that failure mode with a leaderboard attached. If your agent’s “eval” is another LLM’s vibe check, you are compounding the circularity Gienapp is measuring on IR collections.
What this means if you ship search or agents
Most teams still optimize for a number that assumes the labels are settled. New retrieval methods, new agent tools, new context strategies — all scored against pools and judges that may not be the contract you think they are. If unjudged documents are silently non-relevant, you punish systems that find new good documents. If a prompted LLM fills the holes, you may preserve system order while corrupting the labels — lenient, even, and still wrong.
The practical move is narrower than “ban LLM judges.” Use them where you have no pool and no alternative. Prefer human labels where the decision is load-bearing. When you must extend an existing pool, prefer something that calibrates to the assessor you already paid for — topic-specific, small, deterministic, shareable — over a general model that was never shown that assessor’s judgments. Keep the adapter out of the ranker. Keep topic transfer off the table. Keep distillation of adapter labels out of training.
For agent evals, the analogy is direct. If the same model family writes the answer and grades it, you have a circular leaderboard. Separate the assessor. Calibrate when you can. Report uncertainty when you cannot. Fiedler’s sign-reversal is the cautionary slide for anyone who thought shared calibration was free.
What not to claim
This is for extending existing human pools, not for judging from scratch with zero labels. On a brand-new topic, LLM-as-judge may still be the only option — say so, and treat the scores as provisional. Do not claim 26 million parameters beat GPT at understanding relevance. They beat it at matching one assessor. Do not cite Spearman ρ ≥ 0.96 as “the adapter is as good as humans” in every setting; that is ranking correlation at k=50 on the paper’s setup. Do not put the adapter in production retrieval. Do not transfer across topics. Do not distill.
If your eval can also be your ranker, it is not an eval. Calibrate to the human. Then stop.
Papers: Gienapp et al. (SIGIR 2026 Best Student Paper). Support: Otero and Parapar, Lee et al., Fiedler.