Skip to content

Blog

Stop spending the cross-encoder inside the window

When the CE budget is smaller than the context capacity, scoring the obvious top is waste. BoundaryMORPH spends the budget on the k-boundary instead.

Algebraic Retrieval gave the agent a dial for how to score. Production RAG still wastes the expensive dial that decides which documents enter the context. You run a dual-encoder, take a candidate pool, and burn a fixed cross-encoder (CE) budget B on the head of that pool. The generator then gets a context capacity of k documents. On diffuse queries — open-ended research asks with tens or hundreds of relevant sources — B is often smaller than k. Scoring the MIPS top-B and hoping the tail fills the remaining slots is not a budget policy. It is a habit.

Caplan, Roy, Dasgupta, Wang, and Gangadharaiah (Purdue / AWS AI Labs, September 2026) name the habit and replace it (arXiv:2609.27213). BoundaryMORPH treats document selection for RAG as budgeted top-k set membership. The CE calls go where membership is contested — the k-boundary — not where the ranking is already settled.

The mismatch you already ship

Modern RAG is two-stage for a reason. A dual-encoder returns a large candidate pool by maximum inner product search (MIPS). A cross-encoder then scores pairs more accurately, and more slowly, so latency caps CE calls at budget B. Separately, the generator reserves only k slots for retrieved documents. Both limits are known at query time. Most stacks ignore that joint knowledge.

Single-stage rerank scores the MIPS top-B, parks those above the unscored tail, and leaves the tail in MIPS order. When B < k, the remaining slots are filled by documents the CE never saw. On diffuse queries, relevant material is spread through the pool. The paper’s running example: 144 gold-relevant documents, k = 100, B = 25. Twenty-five CE calls cannot verify one hundred slots. Polishing documents already destined for the window buys almost nothing for set quality.

That is the structural flaw. Confirming the obvious top does not change membership. Exploring regions securely below the cutoff does not either. Only calls that could move a document across the inclusion boundary change which set the generator reads.

What BoundaryMORPH actually does

The method is a Gaussian Process (GP) surrogate over the candidate pool, specialized in two places.

First, the prior. Instead of a zero-mean GP that rebuilds relevance from scratch, BoundaryMORPH sets the prior mean to each document’s MIPS percentile rank inside the pool. The GP models the residual between CE score and that prior. Far from any observation the posterior falls back to MIPS — each CE call bends the dual-encoder surface rather than discarding it. A zero-cost observation of the query at maximal relevance anchors the surrogate before the budget starts.

Second, the acquisition rule. Standard GP-UCB seeks a global maximum — the wrong objective when you must return a set of size k. BoundaryMORPH partitions the pool by posterior mean into incumbents (running top-k) and challengers. It picks the contested incumbent–challenger pair with the smallest standardized boundary gap and spends the next CE call on whichever of the two is less certain. That call propagates through the kernel to unscored neighbors, so B observations inform far more than B documents.

After the budget is spent, the returned set is the top-k under the posterior mean. The objective is set relevance (nCG@k), not within-set order: of the k documents the generator can read, how much relevance did you deliver?

What the numbers say

They evaluate on NeuCLIR (English monolingual), Robust04, and TravelDest after filtering to queries with at least twenty relevant documents (85 / 187 / 99 queries; median relevance counts 51 / 62 / 462). The candidate pool is the top N = 10,000 by MIPS. Budgets are B ∈ {25, 50}; capacities are k ∈ {50, 100, 200}. Baselines include BM25 over the shared pool, single-stage CE rerank, RGS (graph expansion), and BAGEL (a prior GP with zero mean and max-seeking UCB).

BoundaryMORPH wins every dataset × budget × capacity cell in Table 1. The abstract headline is +5.4 nCG@100 over the strongest baseline — Robust04 at B = 25, k = 100 (46.4 vs BAGEL’s 41.0). Averaged across the grid, the paper reports +3.8 pp nCG@100 and +7.3 pp CE-nCG@200 over the second-best baseline. Ablations: removing boundary acquisition drops nCG@k by 1.5–2.6 percentage points; removing the MIPS prior drops it by 2.1–3.8. Both pieces earn their keep.

The geometric analysis is the so-what for deep research. Diffuse queries are often multimodal — several semantic peaks rather than one blob. Max-seeking baselines over-measure the primary mode. BoundaryMORPH, by declining to re-score documents safely inside the top-k, keeps budget to cross valleys and pick up secondary peaks. On NeuCLIR, multimodal mass predicts per-query advantage over BAGEL at Spearman ρ ≈ +0.59 (p = 3.5×10⁻⁹). That is the research-agent regime: “effects of light pollution” is not one document; it is a set of facets.

Why production RAG should care

If you run deep research or any RAG path that expects many sources per answer, you already live in the B < k regime more often than design docs admit. CE latency is the expensive dial; context capacity is the scarce slot budget. “Score the top B and dump them plus a MIPS tail” misallocates both.

This sits next to posts already on this site. Stop giving the agent a search box argued the query surface should be a program the agent can revise. BoundaryMORPH is the systems sequel for the rerank dial: once k and B are known, selection is budgeted set membership, not max-seeking polish of the head. Give the agent the retrieval controls argued the interface matters more than the index ritual. Here the interface is the acquisition rule — aim calls at contested membership and return a set sized to the window you reserved.

Agentic loops that treat the retriever as a tool can swap a max-seeking rerank for a boundary selector at equal CE budget. The paper’s latency appendix is blunt: on their NeuCLIR setup, CE inference still dominates wall clock at realistic budgets; the GP bookkeeping is not the bottleneck until B gets extreme.

What not to claim

The method is scoped to diffuse queries (≥20 relevant documents in their filter). For factoid k = 1, max-seeking is fine. CE calls are sequential because the posterior updates after each score — friendly to quality-sensitive deep research, less friendly to heavily batched serving unless you batch the acquisition. Capacity k is a document count, not a token budget; chunked corpora make that a reasonable proxy, not a universal law. Exact GP updates scale as O(B³N); fine at B ≤ 200 in their measurements, painful if CE ever becomes cheap enough that B = 500 is normal. And nCG@k asks about set quality, not within-context ordering, which the authors correctly treat as a separate problem.

Do not read “+5.4” as a portable constant. It is the abstract’s peak delta on this grid; the average nCG@100 lift over the second-best baseline is +3.8 pp. Both are real. Neither replaces measuring your B, k, and query diffuseness.

Build implication

Expose context capacity k and CE budget B as first-class controls on the retrieval path. When B < k and the query is diffuse, stop spending the cross-encoder on documents whose membership is already decided. Aim the budget at the k-boundary, keep a dual-encoder prior so unscored documents still have a calibrated estimate, and score set quality — how much relevance actually entered the window — instead of only how faithfully you polished the head.

Stop spending the cross-encoder inside the window. Spend it where inclusion is still in doubt.