Skip to content

Blog

Your hybrid top-k is a different ranker

Truncating dense and sparse lists before fusion does not approximate complete-list RRF. It computes a different ranking — and the cutoff that worked last quarter may not transfer.

If you ship hybrid RAG, you probably fuse a dense list and a sparse list, then take the top k. The usual way to make that fast is to cap each list at some top_k or rank_window_size — 50, 100, 200 — and treat everything past that cutoff as if it scored zero. Elasticsearch’s rank_window_size and Azure AI Search’s finite text and vector result sets are the production form of this design. Most teams call the result “hybrid RRF” and move on.

That cutoff is not an optimization of the ranker you think you deployed. It is a different ranking function.

I spent years in Solr and OpenSearch land where “window size” was an execution budget you argued about in a runbook. The RAG era quietly turned that budget into the model. Zhang’s August 2026 paper on Exact Adaptive Hybrid Retrieval (EAHR) is the cleanest statement I have seen of why that matters — and what to do instead of pretending a fixed L is portable.

What hybrid RAG actually does

Hybrid retrieval is the default answer to a real problem. Sparse methods (BM25, learned sparse) catch exact terms, identifiers, and rare phrases. Dense methods catch paraphrase and topical similarity. Fuse the two lists and you often get better recall than either alone. Reciprocal rank fusion (RRF) is the workhorse: it scores a document from its ranks in each list, not from raw scores that live on incompatible scales. Weighted RRF and cousins are variants of the same idea.

The production shortcut is truncation. You ask the dense retriever for its top L, the sparse retriever for its top L, fuse those two finite sets, and return the fused top-k. Everything past L is treated as absent. That looks like an approximation of “complete-list” fusion — fuse the full rankings, then cut. It is not.

Truncation changes membership, not just cost

RRF cares about ranks. If a document is rank 4 in dense and unread in sparse, the sparse contribution is missing. The unread rank is not zero. It is unknown. An unknown rank can still move the document into or out of the fused top-k, or reorder the ones already there.

Zhang makes this precise. Even when the truncated candidate union already contains every document that belongs in the complete-list top-20, the order is often wrong. At a window of 100, ordered agreement with complete-list weighted RRF was 44.9%, 41.7%, 16.0%, 44.2%, and 31.5% across the five collections they measured. The documents were in the pool. The ranking was not the ranking.

That is the sentence I want product and infra teams to sit with. If you have ever set top_k=100 on both channels and called the result “hybrid RRF,” that is the system you have. Your design doc may say complete-list fusion. Your serving path computes truncated fusion. Those are different functions. Similar nDCG on a held-out slice does not prove you returned the same list.

The best cutoff does not transfer

Because L decides which rank contributions enter fusion, it is a quality parameter pretending to be an execution budget. It is fit to a query sample and a corpus snapshot. New queries change the dense and sparse rankings. Corpus updates change them again.

On the same MS MARCO index, with the same representations and the same fusion settings, TREC-DL 2019 wanted L = 5000 for highest nDCG@10. TREC-DL 2020 wanted L = 20. Deeper is not always better: on TREC-DL 2020, nDCG@10 peaked at 0.6424 at L = 20, then fell to 0.6179 at L = 100. On TREC-COVID, the curve ran the other way — 0.6624 at L = 10, 0.7925 at L = 100, 0.8059 for the complete list.

A depth chosen on last quarter’s queries is not a contract. It is a guess about this quarter’s rankings. That is why “we tuned rank_window_size once and it was fine” is a fragile story. The cutoff that worked on one query mix can quietly hurt another — same index, opposite “best” depths.

Fix the result, adapt the work

The useful move is to separate the two jobs that top_k currently holds. Declare the retrieval target: the ordered top-k of complete-list fusion. Treat each channel’s depth as internal execution state.

That is the contract behind Exact Adaptive Hybrid Retrieval (EAHR). Dense and sparse retrievers expose resumable exact prefixes — they can produce the next rank without revising earlier ones. Fusion bounds what unread ranks can still contribute, and asks for more only while those bounds can change top-k membership or order. If the lists agree near the top, it stops early. If they are anti-correlated, it may have to exhaust both. Either way, a successful request matches complete-list fusion in items and order.

On Qdrant v1.18.2, EAHR reproduced the complete-list ordered top-20 in all 150 TREC-COVID query–snapshot combinations. Against exhaustive batch execution, geometric-mean latency ratios were about 23× on TREC-DL 2019 and 30× on TREC-DL 2020. That is not a guarantee for every deployment. Anti-correlated rankings exhausted both lists, and some hard queries were slower on the incremental path. The point is not that every request gets cheaper. The point is that the result is the one you named, and the work follows the current rankings instead of a window you picked last month.

If you have lived with OpenSearch or Solr, this should feel familiar. You already distinguish the query you meant from the shard fan-out and windowing that execute it. EAHR is that distinction applied to hybrid fusion: name the ranking function, then adapt the depth.

What to do on Monday

If you run hybrid RAG in production, three checks are worth more than another embedding swap.

Name the ranker. Write down whether the intended target is complete-list RRF or truncated RRF. They are not the same function. If truncated is what you want — latency budget, product choice, known tradeoff — say so, and treat L as part of the model, not a performance knob you can change without re-evaluating quality.

Stop treating L as portable. If you must keep a fixed window, re-select it when the query mix or the corpus moves. The TREC-DL 2019 vs 2020 split is the cautionary slide: same index, opposite “best” depths. A quarterly re-tune of the window is cheaper than a silent ranking drift that your agent then “reasons” over.

Measure ranking agreement, not just nDCG. Similar nDCG does not mean you returned the same list. Zhang’s L = 100 windows already had the right documents in the union and still lost ordered agreement. If you A/B a cutoff, look at membership and order against a complete-list (or much deeper) baseline on a held-out slice. Agreement is the metric that matches the claim you are making about the ranker.

What not to claim

Do not read 23× / 30× as a universal latency win for every hybrid stack. Those numbers are against exhaustive batch execution on specific collections and a specific Qdrant build. Anti-correlated lists and hard queries can erase the savings. Do not install EAHR as a drop-in for every Elasticsearch rank_window_size without checking whether your dense and sparse backends can expose resumable exact prefixes. And do not treat ordered agreement as the only product metric — users care about nDCG and downstream answer quality too. The point is that those metrics are meaningless as a comparison to “RRF” if you never named which RRF you meant.

Context engineering is, at its core, a search problem. The context window only ever sees what retrieval put there. If the fusion step silently computes a different ranking than the one in the design doc, the agent is not failing at reasoning. It is reading the wrong evidence.

Paper: Exact Adaptive Hybrid Retrieval Without Fixed Top-L Cutoffs (Zhang, 2026). Companion on exact dense search: DESA.