Don't decode a ranking as prose
Generative listwise rerankers won on quality. Autoregressive decode still bills you per token. SPD reads the ranking off the prefill and solves the assignment instead.
If you run a serious retrieval stack, you already know the last stage is where quality and latency argue. Dense and sparse first-stage retrieval give you a candidate set. A reranker is supposed to put that set in the right order before the generator or the UI sees it. Pointwise cross-encoders still dominate production because they are fast and boring. Listwise LLM rerankers — RankGPT and its descendants — are often better because they see the whole slate and emit a joint ordering. They are also slow, because “emit a joint ordering” usually means left-to-right token generation.
That is the wrong shape for the problem.
Laftchiev, Agrawal, Kayali, Yan, Xu, Lei, Qiu, Hua, Li, and Simon (Meta, September 2026) call the method SPD: Single Pass Decoding for generative reranking (arXiv:2609.01807). The claim is narrow and sharp. The only tokens a ranker must emit are the N ordinals that name the items in ranked order. That output is a permutation, not open-ended prose. Once you admit that, you can stop pretending ranking is chat and decode it as combinatorial assignment instead of autoregressive text.
What generative listwise ranking actually costs
A listwise LLM reranker takes a query (or user context) and N candidates in one prompt. Prefill runs once over that whole prompt. Then the model generates the ranking as tokens — item identifiers or ordinals, sometimes with a reasoning trace first. Each emitted token is another sequential forward pass over the KV cache. Prefill is already parallel. Decode is not.
The paper’s teacher numbers make the split concrete. On their proprietary reranking set, an autoregressive teacher with an explicit reasoning trace takes about 1807 ms per request on an A100. Of that, roughly 28 ms is prefill. Strip the reasoning and keep only the ranking tokens and you still sit around 88 ms. The comparative information that is the ranking was already computed during the parallel pass over the candidates. Autoregressive decoding re-serializes it.
That is why many production teams never ship the generative listwise stage even when offline evals look good. Latency budgets for search and ads are measured in tens of milliseconds, not seconds. Pointwise models win the bake-off on serving cost. Quality stays on the whiteboard.
What SPD does instead
SPD keeps the same intended output — the N ordinals — and changes only how they are decoded. After one prefill over all candidates, a light self-attention head reads an N × K item–position score matrix off the backbone’s hidden states. The Hungarian algorithm (or LAPJV) solves maximum-weight bipartite matching on that matrix. The result is a valid permutation by construction: every item gets exactly one rank, every rank gets exactly one item. No duplicate-rank repair, no sampling temperature, no beam-search artifacts.
Training is distillation, not magic. An autoregressive teacher labels permutations offline. The student — a 0.6B decoder-only backbone with LoRA, plus the scoring head — learns to reproduce those permutations via a Sinkhorn-relaxed objective at train time, then switches to hard Hungarian assignment at inference. The interesting ablation is not “does distillation help?” It is that teacher permutations alone are not enough on a frozen backbone. Click labels on a frozen model are weak. Full permutation targets on a frozen model are still weak. Only teacher rankings plus LoRA adaptation close the capacity gap so the hidden states actually encode pairwise preferences the assignment can use.
On the internal set: SPD lands at 28 ms end-to-end — about 64× the reasoning teacher and about 3× the no-reasoning teacher that already emits only ranking tokens. List-level metrics stay on par (NDCG@1 and Recall@1 indistinguishable; AUC retains essentially all of the teacher). Latency breakdown is almost entirely backbone prefill (~27.9 ms). The scoring head is under 0.1 ms. Assignment for N = 50 is about 0.008 ms on CPU. The combinatorial solver is not the tax people fear. The sequential decode was.
They also run Amazon Beauty with a Qwen3-32B teacher and a 0.6B SPD student: roughly 45× faster than the 32B teacher, quality close to it and far above a plain 0.6B autoregressive baseline. Treat that as corroboration on a public set, not as a claim that every domain will look identical.
Why this matters if you ship search
For teams that already run hybrid retrieval plus a reranker — the stack we have been writing about all month — the question has quietly shifted. It is no longer “can an LLM rank a slate?” RankGPT answered that years ago. It is “can we afford the decode inside the latency envelope where search and ads live?”
SPD’s answer is: yes, if you stop generating the ranking as language. Ranking has a known feasible set. The output alphabet is {1…N}, the length is known before decode starts, and every value must appear once. That is an assignment problem with a classical exact solver, not a chat completion. Framing generative ranking as combinatorial optimization is the systems punchline. It also explains a quiet advantage over the teacher: the Hungarian step cannot emit an invalid ordering. Autoregressive models can and do.
This sits next to the other threads on this blog without repeating them. Hybrid top-k truncation changes which ranker you actually deployed. Hybrid search as a program argues fusion itself should be executable, not a fixed weight. SPD is the complementary claim at the listwise stage: the decoder should respect the structure of the object you are producing. If the object is a permutation, decode a permutation.
There is a product implication worth stating plainly. Speculative decoding, multi-token heads, and shorter prompts are useful accelerations of left-to-right generation. They still treat ranking as text that happens to be short. SPD refuses that framing. Once decode cost collapses to prefill, further latency wins have to come from the backbone — quantization, pruning, early exit — not from cleverer token schedules. That is a cleaner roadmap for infra teams than another year of “make the chat model emit ranks faster.”
What not to claim
Do not install 28 ms or 64× as portable constants. The headline numbers are on a proprietary Meta reranking set with a matched 0.6B teacher/student pair on A100 hardware. Amazon Beauty corroborates the speedup pattern; it does not make the absolute milliseconds universal. Hungarian / LAPJV is fine at typical rerank depths (N ≈ 10–50). It is not an invitation to dump a thousand-item slate into one assignment solve and call it free. Distillation quality is capped by the teacher — SPD transfers a ranking, it does not invent a better one than the model that labeled the data. And this paper is about the rerank stage after an upstream retriever already proposed candidates. It does not replace first-stage retrieval, hybrid fusion, or grounding gates.
Cite it as SPD / 2609.01807. Some secondary indexes still show an earlier “hLLM” label; the arXiv abs title is SPD.
Build implication
If you are evaluating LLM listwise rerankers for production search, stop benchmarking only offline nDCG of an autoregressive decode. Benchmark the serving path: p95 latency at your real slate size, validity rate of emitted permutations, and quality after you have stripped every token that is not an ordinal. If quality collapses when you remove the reasoning trace, you did not have a ranker — you had a chatty judge. If quality holds and latency does not, you have SPD’s problem statement. Then ask whether your decoder should still be left-to-right language modeling at all.
Generative listwise ranking won on quality. Decode still bills you per token. The ranking was already in the prefill. If the output is a permutation, don’t pretend it’s prose.